A survey of large-scale analytical query processing in MapReduce

Christos Doulkeridis, Kjetil Nørvåg
<span title="2013-06-08">2013</span> <i title="Springer Nature America, Inc"> <a target="_blank" rel="noopener" href="https://fatcat.wiki/container/bj2uhzcrunglzmjsyjjm3qh4hm" style="color: black;">The VLDB journal</a> </i> &nbsp;
Enterprises today acquire vast volumes of data from different sources and leverage this information by means of data analysis to support effective decision-making and provide new functionality and services. The key requirement of data analytics is scalability, simply due to the immense volume of data that need to be extracted, processed, and analyzed in a timely fashion. Arguably the most popular framework for contemporary large-scale data analytics is MapReduce, mainly due to its salient
more &raquo; ... es that include scalability, fault-tolerance, ease of programming, and flexibility. However, despite its merits, MapReduce has evident performance limitations in miscellaneous analytical tasks, and this has given rise to a significant body of research that aim at improving its efficiency, while maintaining its desirable properties. This survey aims to review the state of the art in improving the performance of parallel query processing using MapReduce. A set of the most significant weaknesses and limitations of MapReduce is discussed at a high level, along with solving techniques. A taxonomy is presented for categorizing existing research on MapReduce improvements according to the specific problem they target. Based on the proposed taxonomy, a classification of existing research is provided focusing on the optimization objective. Concluding, we outline interesting directions for future parallel data processing systems.
<span class="external-identifiers"> <a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.1007/s00778-013-0319-9">doi:10.1007/s00778-013-0319-9</a> <a target="_blank" rel="external noopener" href="https://fatcat.wiki/release/3gkpguiwnre2jduhjssuqgydfq">fatcat:3gkpguiwnre2jduhjssuqgydfq</a> </span>
<a target="_blank" rel="noopener" href="https://web.archive.org/web/20170808031516/http://romisatriawahono.net/lecture/rm/survey/information%20retrieval/Doulkeridis%20-%20large-scale%20analytical%20query%20processing%20in%20MapReduce%20-%202014.pdf" title="fulltext PDF download" data-goatcounter-click="serp-fulltext" data-goatcounter-title="serp-fulltext"> <button class="ui simple right pointing dropdown compact black labeled icon button serp-button"> <i class="icon ia-icon"></i> Web Archive [PDF] <div class="menu fulltext-thumbnail"> <img src="https://blobs.fatcat.wiki/thumbnail/pdf/fb/ca/fbcac4f049331354ada4cca50e6fba53d120338b.180px.jpg" alt="fulltext thumbnail" loading="lazy"> </div> </button> </a> <a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.1007/s00778-013-0319-9"> <button class="ui left aligned compact blue labeled icon button serp-button"> <i class="external alternate icon"></i> springer.com </button> </a>