Regularized Latent Semantic Indexing-Reference-Cited by-同舟云学术

Regularized Latent Semantic Indexing

Published:2013-01 Issue:1 Volume:31 Page:1-44
ISSN:1046-8188
Container-title:ACM Transactions on Information Systems
language:en
Short-container-title:ACM Trans. Inf. Syst.

Author:

Wang Quan¹,Xu Jun²,Li Hang²,Craswell Nick³

Affiliation:

1. MOE-Microsoft Key Laboratory of Statistics and Information Technology of Peking University

2. Microsoft Research Asia

3. Microsoft Corporation

Abstract

Topic modeling provides a powerful way to analyze the content of a collection of documents. It has become a popular tool in many research areas, such as text mining, information retrieval, natural language processing, and other related fields. In real-world applications, however, the usefulness of topic modeling is limited due to scalability issues. Scaling to larger document collections via parallelization is an active area of research, but most solutions require drastic steps, such as vastly reducing input vocabulary. In this article we introduce Regularized Latent Semantic Indexing (RLSI)---including a batch version and an online version, referred to as batch RLSI and online RLSI, respectively---to scale up topic modeling. Batch RLSI and online RLSI are as effective as existing topic modeling techniques and can scale to larger datasets without reducing input vocabulary. Moreover, online RLSI can be applied to stream data and can capture the dynamic evolution of topics. Both versions of RLSI formalize topic modeling as a problem of minimizing a quadratic loss function regularized by ℓ1 and/or ℓ2 norm. This formulation allows the learning process to be decomposed into multiple suboptimization problems which can be optimized in parallel, for example, via MapReduce. We particularly propose adopting ℓ1 norm on topics and ℓ2 norm on document representations to create a model with compact and readable topics and which is useful for retrieval. In learning, batch RLSI processes all the documents in the collection as a whole, while online RLSI processes the documents in the collection one by one. We also prove the convergence of the learning of online RLSI. Relevance ranking experiments on three TREC datasets show that batch RLSI and online RLSI perform better than LSI, PLSI, LDA, and NMF, and the improvements are sometimes statistically significant. Experiments on a Web dataset containing about 1.6 million documents and 7 million terms, demonstrate a similar boost in performance.

Publisher

Association for Computing Machinery (ACM)

Subject

Computer Science Applications,General Business, Management and Accounting,Information Systems

Link

https://dl.acm.org/doi/pdf/10.1145/2414782.2414787

Reference57 articles.

1. On-line LDA: Adaptive Topic Models for Mining Text Streams with Applications to Topic Detection and Tracking

2. Asuncion A. Smyth P. and Welling M. 2011. Asynchronous distributed estimation of topic models for document analysis. Stat. Methodol. Asuncion A. Smyth P. and Welling M. 2011. Asynchronous distributed estimation of topic models for document analysis. Stat. Methodol .

3. Latent semantic indexing (LSI) fails for TREC collections

4. Bertsekas D. P. 1999. Nonlinear Programming. Athena Scientific Belmont MA. Bertsekas D. P. 1999. Nonlinear Programming . Athena Scientific Belmont MA.

Cited by 31 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. A Summary of Unsupervised Learning Methods;Machine Learning Methods;2023-12-06

2. Influence of the Spatial Distribution of Jobs in Intervening Opportunities Models;Transportation Research Record: Journal of the Transportation Research Board;2023-01-06

3. Improved Evolutionary Approach for Tuning Topic Models with Additive Regularization;Lecture Notes in Computer Science;2023

4. Topic Modelling for Research Perception: Techniques, Processes and a Case Study;Recent Innovations in Artificial Intelligence and Smart Applications;2022

5. Fine-Grained Privacy Detection with Graph-Regularized Hierarchical Attentive Representation Learning;ACM Transactions on Information Systems;2020-10-13