Ultra-high throughput sequencing-based small RNA discovery and discrete statistical biomarker analysis in a collection of cervical tumours and matched controls

Daniela Witten, Robert Tibshirani, Sam Gu, Andrew Fire, Weng-Onn Lui
2010 BMC Biology  
Ultra-high throughput sequencing technologies provide opportunities both for discovery of novel molecular species and for detailed comparisons of gene expression patterns. Small RNA populations are particularly well suited to this analysis, as many different small RNAs can be completely sequenced in a single instrument run. Results: We prepared small RNA libraries from 29 tumour/normal pairs of human cervical tissue samples. Analysis of the resulting sequences (42 million in total) defined 64
more » ... w human microRNA (miRNA) genes. Both arms of the hairpin precursor were observed in twenty-three of the newly identified miRNA candidates. We tested several computational approaches for the analysis of class differences between high throughput sequencing datasets and describe a novel application of a log linear model that has provided the most effective analysis for this data. This method resulted in the identification of 67 miRNAs that were differentially-expressed between the tumour and normal samples at a false discovery rate less than 0.001. Conclusions: This approach can potentially be applied to any kind of RNA sequencing data for analysing differential sequence representation between biological sample sets. Cite this article as: Witten et al., Ultra-high throughput sequencing-based small RNA discovery and discrete statistical biomarker analysis in a collection of cervical tumours and matched controls BMC Biology 2010, 8:58
doi:10.1186/1741-7007-8-58 pmid:20459774 pmcid:PMC2880020 fatcat:ifslgcjjgneh7da62qg74ho3d4