Incremental and Parallel Analytics on Astrophysical Data Streams

Dmitry Mishin, Tamas Budavari, Alexander Szalay, Yanif Ahmad
2012 2012 SC Companion: High Performance Computing, Networking Storage and Analysis  
Stream processing methods and online algorithms are increasingly appealing in the scientific and large-scale data management communities due to increasing ingestion rates of scientific instruments, the ability to produce and inspect results interactively, and the simplicity and efficiency of sequential storage access over enormous datasets. This article will showcase our experiences in using off-the-shelf streaming technology to implement incremental and parallel spectral analysis of galaxies
more » ... om the Sloan Digital Sky Survey (SDSS) to detect a wide variety of galaxy features. The technical focus of the article is on a robust, highly scalable principal components analysis (PCA) algorithm and its use of coordination primitives to realize consistency as part of parallel execution. Our algorithm and framework can be readily used in other domains.
doi:10.1109/sc.companion.2012.130 dblp:conf/sc/MishinBSA12 fatcat:mcw2iozawbfkre2woggedka2ue