Scalable topical phrase mining from text corpora

Ahmed El-Kishky, Yanglei Song, Chi Wang, Clare R. Voss, Jiawei Han
2014 Proceedings of the VLDB Endowment  
While most topic modeling algorithms model text corpora with unigrams, human interpretation often relies on inherent grouping of terms into phrases. As such, we consider the problem of discovering topical phrases of mixed lengths. Existing work either performs post processing to the results of unigram-based topic models, or utilizes complex n-gramdiscovery topic models. These methods generally produce low-quality topical phrases or suffer from poor scalability on even moderately-sized datasets.
more » ... We propose a different approach that is both computationally efficient and effective. Our solution combines a novel phrase mining framework to segment a document into single and multi-word phrases, and a new topic model that operates on the induced document partition. Our approach discovers high quality topical phrases with negligible extra cost to the bag-of-words topic model in a variety of datasets including research publication titles, abstracts, reviews, and news articles.
doi:10.14778/2735508.2735519 fatcat:l5dgrmk3sngg3aghbnyrbahmli