Sense Cluster Based Categorization and Clustering of Abstracts [chapter]

Davide Buscaldi, Paolo Rosso, Mikhail Alexandrov, Alfons Juan Ciscar
2006 Lecture Notes in Computer Science  
This paper focuses on the use of sense clusters for classification and clustering of very short texts such as conference abstracts. Common keyword-based techniques are effective for very short documents only when the data pertain to different domains. In the case of conference abstracts, all the documents are from a narrow domain (i.e., share a similar terminology), that increases the difficulty of the task. Sense clusters are extracted from abstracts, exploiting the WordNet relationships
more » ... ng between words in the same text. Experiments were carried out both for the categorization task, using Bernoulli mixtures for binary data, and the clustering task, by means of Stein's MajorClust method.
doi:10.1007/11671299_56 fatcat:laxd2qcggff5rcdtwuqfu4vcfy