Variable selection in model-based clustering and discriminant analysis with a regularization approach [article]

Gilles Celeux, Cathy Maugis-Rabusseau, Mohammed Sedki
2017 arXiv   pre-print
Relevant methods of variable selection have been proposed in model-based clustering and classification. These methods are making use of backward or forward procedures to define the roles of the variables. Unfortunately, these stepwise procedures are terribly slow and make these variable selection algorithms inefficient to treat large data sets. In this paper, an alternative regularization approach of variable selection is proposed for model-based clustering and classification. In this approach,
more » ... the variables are first ranked with a lasso-like procedure in order to avoid painfully slow stepwise algorithms. Thus, the variable selection methodology of Maugis et al (2009b) can be efficiently applied on high-dimensional data sets.
arXiv:1705.00946v1 fatcat:k7mv4p6zmzderk27cp4atduupq