Towards a Cloud Native Big Data Platform using MiCADO

Abdelkhalik Mosa, Tamas Kiss, Gabriele Pierantoni, James DesLauriers, Dimitrios Kagialis, Gabor Terstyanszky
2020 2020 19th International Symposium on Parallel and Distributed Computing (ISPDC)  
obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. The WestminsterResearch online digital archive at the University of Westminster aims to make the research output of the University available to a wider audience. Abstract-In the big data era, creating
more » ... elf-managing scalable platforms for running big data applications is a fundamental task. Such self-managing and self-healing platforms involve a proper reaction to hardware (e.g., cluster nodes) and software (e.g., big data tools) failures, besides a dynamic resizing of the allocated resources based on overload and underload situations and scaling policies. The distributed and stateful nature of big data platforms (e.g., Hadoop-based cluster) makes the management of these platforms a challenging task. This paper aims to design and implement a scalable cloud native Hadoopbased big data platform using MiCADO, an open-source, and a highly customisable multi-cloud orchestration and auto-scaling framework for Docker containers, orchestrated by Kubernetes. The proposed MiCADO-based big data platform automates the deployment and enables an automatic horizontal scaling (in and out) of the underlying cloud infrastructure. The empirical evaluation of the MiCADO-based big data platform demonstrates how easy, efficient, and fast it is to deploy and undeploy Hadoop clusters of different sizes. Additionally, it shows how the platform can automatically be scaled based on user-defined policies (such as CPU-based scaling).
doi:10.1109/ispdc51135.2020.00025 dblp:conf/ispdc/MosaKPDKT20 fatcat:cyxdwx7gjfgu3lpng5ueyahcly