A Cooperative Binary-Clustering Framework Based on Majority Voting for Twitter Sentiment Analysis

Maryum Bibi, Wajid Aziz, Majid Almarashi, Imtiaz Hussain Khan, Malik Sajjad Ahmed Nadeem, Nazneen Habib
2020 IEEE Access  
Twitter sentiment analysis is a challenging problem in natural language processing. For this purpose, supervised learning techniques have mostly been employed, which require labeled data for training. However, it is very time consuming to label datasets of large size. To address this issue, unsupervised learning techniques such as clustering can be used. In this study, we explore the possibility of using hierarchical clustering for twitter sentiment analysis. Three hierarchical-clustering
more » ... ques, namely single linkage (SL), complete linkage (CL) and average linkage (AL), are examined. A cooperative framework of SL, CL and AL is built to select the optimal cluster for tweets wherein the notion of optimal-cluster selection is operationalized using majority voting. The hierarchical clustering techniques are also compared with k-means and two state-of-the-art classifiers (SVM and Naïve Bayes). The performance of clustering and classification is measured in terms of accuracy and time efficiency. The experimental results indicate that cooperative clustering based on majority voting approach is robust in terms of good quality clusters with tradeoff of poor time efficiency. The results also suggest that the accuracy of the proposed clustering framework is comparable to classifiers which is encouraging. INDEX TERMS Cooperative clustering, majority voting, sentiment analysis, twitter sentiment analysis.
doi:10.1109/access.2020.2983859 fatcat:yazthauwr5fzvifbwmq74jrsqu