Efficient Software for Archiving and Retrieving Results of Massive Bioinformatics Analyses in High-Performance Computing Environments

Craig P Steffen, Roland Haas, Katherine Kendig, Liudmila Mainzer, Ryan Chui, Christina Fliege
<span title="2021-12-27">2021</span> <i title="Zenodo"> Zenodo </i> &nbsp;
Abstract—Modern sequencing and computational facilities in biomedical and agricultural areas generate and analyze hundreds or thousands of samples every day. At that scale production bioin- formatics workflows can produce vast amounts of data, which need to be managed: organized, deleted, moved, stored, and retrieved. This can create an infrastructure bottleneck especially when moving or archiving data. The difficulty stems chiefly from the structure of file collections being produced:
more &raquo; ... y very large number of files, with a highly nested directory structure, and a heterogeneous distribution of file sizes, with emphasis on large numbers of very small files. Parallel file systems, such as Lustre, GPFS, and tape archives, can perform poorly under these circumstances due to overabundance of metadata. However, standard packaging utilities, such as tar and zip, do not scale well with the size of the data for this particular use case. The present manuscript reviews several recently developed parallel alternatives, showcasing their performance on a variety of high performance computing systems.
<span class="external-identifiers"> <a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.5281/zenodo.5805629">doi:10.5281/zenodo.5805629</a> <a target="_blank" rel="external noopener" href="https://fatcat.wiki/release/avnfbxcjzvew3p54dui3yxzz7i">fatcat:avnfbxcjzvew3p54dui3yxzz7i</a> </span>
<a target="_blank" rel="noopener" href="https://web.archive.org/web/20220105224501/https://zenodo.org/record/5805629/files/file_packaging_paper_HUST21_2021dec27a.pdf" title="fulltext PDF download" data-goatcounter-click="serp-fulltext" data-goatcounter-title="serp-fulltext"> <button class="ui simple right pointing dropdown compact black labeled icon button serp-button"> <i class="icon ia-icon"></i> Web Archive [PDF] <div class="menu fulltext-thumbnail"> <img src="https://blobs.fatcat.wiki/thumbnail/pdf/72/1c/721c3c45a546a86c62f0c1db991fd39d5fc42cc4.180px.jpg" alt="fulltext thumbnail" loading="lazy"> </div> </button> </a> <a target="_blank" rel="external noopener noreferrer" href="https://doi.org/10.5281/zenodo.5805629"> <button class="ui left aligned compact blue labeled icon button serp-button"> <i class="unlock alternate icon" style="background-color: #fb971f;"></i> zenodo.org </button> </a>