A Multi-Criteria Document Clustering Method Based on Topic Modeling and Pseudoclosure Function

Authors

  • Quang Vu Bui Hue University of Sciences, Vietnam and CHArt Laboratory EA 4004, EPHE
  • Karim Sayadi Sorbonne University, UPMC Univ Paris 06 and CHArt Laboratory EA 4004, EPHE
  • Marc Bui EPHE and UP8 Univ Paris 08 and CHArt Laboratory EA 4004

Abstract

We address in this work the problem of document clustering. Our contribution proposes a novel unsupervised clustering method based on the structural analysis of the latent semantic space. Each document in the space is a vector of probabilities that represents a distribution of topics. The document membership to a cluster is computed taking into account two criteria: the major topic in the document (qualitative criterion) and the distance measure between the vectors of probabilities (quantitative criterion). We perform a structural analysis on the latent semantic space using the Pretopology theory that allows us to investigate the role of the number of clusters and the chosen centroids, in the similarity between the computed clusters. We have applied our method to Twitter data and showed the accuracy of our results compared to a random choice number of clusters. 

Downloads

Published

2016-07-11

How to Cite

Bui, Q. V., Sayadi, K., & Bui, M. (2016). A Multi-Criteria Document Clustering Method Based on Topic Modeling and Pseudoclosure Function. Informatica, 40(2). Retrieved from https://puffbird.ijs.si/index.php/informatica/article/view/1278

Issue

Section

Special issue papers