Efficient parallel spectral clustering algorithm design for large data sets under cloud computing environment

Ran Jin, Chunhai Kou, Ruijuan Liu, Yefeng Li
2013 Journal of Cloud Computing: Advances, Systems and Applications  
Spectral clustering algorithm has proved be more effective than most traditional algorithms in finding clusters. However, its high computational complexity limits its effect in actual application. This paper combines the spectral clustering with MapReduce, through evaluation of sparse matrix eigenvalue and computation of distributed cluster, puts forward the improvement ideas and concrete realization, and thus improves the clustering speed of the distinctive clustering algorithm. According to
more » ... e experiment, with the processing data scale being enlarged, the clustering rate is in nearly linear growth, and the proposed parallel spectral clustering algorithm is suitable for large data mining. The research results provide research basis to better design a clustering partition algorithm in large data and high efficiency.
doi:10.1186/2192-113x-2-18 fatcat:q7gbchi7nrg5rfibzihtzeyhnq