Home LiteratureArticle Details
PMID: 21113236 Published · ppublish English Journal Article

Statistical Analyses of Next Generation Sequence Data: A Partial Overview.

Journal of proteomics & bioinformatics ·Vol. 3 ·No. 6 ·2010-06-01 ·Pages 183-190

Datta S, Datta S, Kim S, Chakraborty S, Gill RS

Abstract

Next generation sequencing has revolutionized the status of biological research. For a long time, the gold standard of DNA sequencing was considered to be the Sanger method. However, in 2005, commercial launching of next generation sequencing has made it possible to generate massively parallel and high resolution DNA sequence data. Its usefulness in various genomic applications such as genome-wide detection of SNPs, DNA methylation profiling, mRNA expression profiling, whole-genome re-sequencing and so on are now well recognized. There are several platforms for generating next generation sequencing (NGS) data which we briefly discuss in this mini overview. With new technologies come new challenges for the data analysts. This mini review attempts to present a collection of selected topics in the current development of statistical methods dealing with these novel data types. We believe that knowing the advances and bottlenecks of this technology will help the researchers to benchmark the analytical tools dealing with these data and will pave the path for its proper application into clinical diagnostics.

Authors & Affiliations
5 authors, click to expand affiliations / ORCID
Datta Susmita
Department of Bioinformatics and Biostatistics, University of Louisville, Louisville, KY 40202, USA.
Datta Somnath
Kim Seongho
Chakraborty Sutirtha
Gill Ryan S
References (66)
66 references, click to expand
  1. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  2. F-Seq: a feature density estimator for high-throughput sequence tags.
    Bioinformatics. 2008 Nov 1;24(21):2537-8 PMID: 18784119
  3. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
  4. Discovering microRNAs from deep sequencing data using miRDeep.
    Nat Biotechnol. 2008 Apr;26(4):407-15 PMID: 18392026
  5. Pyrosequencing applications.
    Methods Mol Biol. 2007;373:15-24 PMID: 17185754
  6. Computational methods for discovering structural variation with next-generation sequencing.
    Nat Methods. 2009 Nov;6(11 Suppl):S13-20 PMID: 19844226
  7. Single-molecule DNA sequencing of a viral genome.
    Science. 2008 Apr 4;320(5872):106-9 PMID: 18388294
  8. Swift: primary data analysis for the Illumina Solexa sequencing platform.
    Bioinformatics. 2009 Sep 1;25(17):2194-9 PMID: 19549630
  9. Hierarchical hidden Markov model with application to joint analysis of ChIP-chip and ChIP-seq data.
    Bioinformatics. 2009 Jul 15;25(14):1715-21 PMID: 19447789
  10. CNV-seq, a new method to detect copy number variation using high-throughput sequencing.
    BMC Bioinformatics. 2009 Mar 06;10:80 PMID: 19267900
  11. Sequence information can be obtained from single DNA molecules.
    Proc Natl Acad Sci U S A. 2003 Apr 1;100(7):3960-4 PMID: 12651960
  12. Transcript length bias in RNA-seq data confounds systems biology.
    Biol Direct. 2009 Apr 16;4:14 PMID: 19371405
  13. An integrated software system for analyzing ChIP-chip and ChIP-seq data.
    Nat Biotechnol. 2008 Nov;26(11):1293-300 PMID: 18978777
  14. Evaluation of statistical methods for normalization and differential expression in mRNA-Seq experiments.
    BMC Bioinformatics. 2010 Feb 18;11:94 PMID: 20167110
  15. Statistical inferences for isoform expression in RNA-Seq.
    Bioinformatics. 2009 Apr 15;25(8):1026-32 PMID: 19244387
  16. Probabilistic base calling of Solexa sequencing data.
    BMC Bioinformatics. 2008 Oct 13;9:431 PMID: 18851737
  17. Improved base calling for the Illumina Genome Analyzer using machine learning strategies.
    Genome Biol. 2009;10(8):R83 PMID: 19682367
  18. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  19. Optimization by simulated annealing.
    Science. 1983 May 13;220(4598):671-80 PMID: 17813860
  20. Statistical model for whole genome sequencing and its application to minimally invasive diagnosis of fetal genetic disease.
    Bioinformatics. 2009 May 15;25(10):1244-50 PMID: 19307238
  21. Sequencing technologies - the next generation.
    Nat Rev Genet. 2010 Jan;11(1):31-46 PMID: 19997069
  22. Transforming single DNA molecules into fluorescent magnetic particles for detection and enumeration of genetic variations.
    Proc Natl Acad Sci U S A. 2003 Jul 22;100(15):8817-22 PMID: 12857956
  23. Next-generation DNA sequencing methods.
    Annu Rev Genomics Hum Genet. 2008;9:387-402 PMID: 18576944
  24. rtracklayer: an R package for interfacing with genome browsers.
    Bioinformatics. 2009 Jul 15;25(14):1841-2 PMID: 19468054
  25. Application of massively parallel sequencing to microRNA profiling and discovery in human embryonic stem cells.
    Genome Res. 2008 Apr;18(4):610-21 PMID: 18285502
  26. Next-generation DNA sequencing.
    Nat Biotechnol. 2008 Oct;26(10):1135-45 PMID: 18846087
  27. A feature-based approach to modeling protein-DNA interactions.
    PLoS Comput Biol. 2008 Aug 22;4(8):e1000154 PMID: 18725950
  28. Design and analysis of ChIP-seq experiments for DNA-binding proteins.
    Nat Biotechnol. 2008 Dec;26(12):1351-9 PMID: 19029915
  29. PeakSeq enables systematic scoring of ChIP-seq experiments relative to controls.
    Nat Biotechnol. 2009 Jan;27(1):66-75 PMID: 19122651
  30. Applications of ultra-high-throughput sequencing.
    Methods Mol Biol. 2009;553:79-108 PMID: 19588102
  31. Four-color DNA sequencing by synthesis using cleavable fluorescent nucleotide reversible terminators.
    Proc Natl Acad Sci U S A. 2006 Dec 26;103(52):19635-40 PMID: 17170132
  32. Stem cell transcriptome profiling via massive-scale mRNA sequencing.
    Nat Methods. 2008 Jul;5(7):613-9 PMID: 18516046
  33. Comparison of next generation sequencing technologies for transcriptome characterization.
    BMC Genomics. 2009 Aug 01;10:347 PMID: 19646272
  34. SNVMix: predicting single nucleotide variants from next-generation sequencing of tumors.
    Bioinformatics. 2010 Mar 15;26(6):730-6 PMID: 20130035
  35. RNA-seq: an assessment of technical reproducibility and comparison with gene expression arrays.
    Genome Res. 2008 Sep;18(9):1509-17 PMID: 18550803
  36. Alta-Cyclic: a self-optimizing base caller for next-generation sequencing.
    Nat Methods. 2008 Aug;5(8):679-82 PMID: 18604217
  37. GenomeGraphs: integrated genomic data visualization with R.
    BMC Bioinformatics. 2009 Jan 06;10:2 PMID: 19123956
  38. The transcriptional landscape of the yeast genome defined by RNA sequencing.
    Science. 2008 Jun 6;320(5881):1344-9 PMID: 18451266
  39. ShortRead: a bioconductor package for input, quality assessment and exploration of high-throughput sequence data.
    Bioinformatics. 2009 Oct 1;25(19):2607-8 PMID: 19654119
  40. Shotgun bisulphite sequencing of the Arabidopsis genome reveals DNA methylation patterning.
    Nature. 2008 Mar 13;452(7184):215-9 PMID: 18278030
  41. Visualization of genomic data with the Hilbert curve.
    Bioinformatics. 2009 May 15;25(10):1231-5 PMID: 19297348
  42. Genome sequencing in microfabricated high-density picolitre reactors.
    Nature. 2005 Sep 15;437(7057):376-80 PMID: 16056220
  43. Model-based analysis of ChIP-Seq (MACS).
    Genome Biol. 2008;9(9):R137 PMID: 18798982
  44. Estimating accuracy of RNA-Seq and microarrays with proteomics.
    BMC Genomics. 2009 Apr 16;10:161 PMID: 19371429
  45. Mapping and quantifying mammalian transcriptomes by RNA-Seq.
    Nat Methods. 2008 Jul;5(7):621-8 PMID: 18516045
  46. Mapping short DNA sequencing reads and calling variants using mapping quality scores.
    Genome Res. 2008 Nov;18(11):1851-8 PMID: 18714091
  47. Annotation of metagenome short reads using proxygenes.
    Bioinformatics. 2008 Aug 15;24(16):i7-13 PMID: 18689842
  48. A clustering approach for identification of enriched domains from histone modification ChIP-Seq data.
    Bioinformatics. 2009 Aug 1;25(15):1952-8 PMID: 19505939
  49. Genome-wide approaches to studying chromatin modifications.
    Nat Rev Genet. 2008 Mar;9(3):179-91 PMID: 18250624
  50. TileQC: a system for tile-based quality control of Solexa data.
    BMC Bioinformatics. 2008 May 28;9:250 PMID: 18507856
  51. DNA sequencing with chain-terminating inhibitors.
    Proc Natl Acad Sci U S A. 1977 Dec;74(12):5463-7 PMID: 271968
  52. Computation for ChIP-seq and RNA-seq studies.
    Nat Methods. 2009 Nov;6(11 Suppl):S22-32 PMID: 19844228
  53. BayesCall: A model-based base-calling algorithm for high-throughput short-read sequencing.
    Genome Res. 2009 Oct;19(10):1884-95 PMID: 19661376
  54. Genome-wide mapping of in vivo protein-DNA interactions.
    Science. 2007 Jun 8;316(5830):1497-502 PMID: 17540862
  55. Modeling ChIP sequencing in silico with applications.
    PLoS Comput Biol. 2008 Aug 22;4(8):e1000158 PMID: 18725927
  56. FindPeaks 3.1: a tool for identifying areas of enrichment from massively parallel short-read sequencing technology.
    Bioinformatics. 2008 Aug 1;24(15):1729-30 PMID: 18599518
  57. Model-based quality assessment and base-calling for second-generation sequencing data.
    Biometrics. 2010 Sep;66(3):665-74 PMID: 19912177
  58. Pyrobayes: an improved base caller for SNP discovery in pyrosequences.
    Nat Methods. 2008 Feb;5(2):179-81 PMID: 18193056
  59. Targeted gene inactivation in zebrafish using engineered zinc-finger nucleases.
    Nat Biotechnol. 2008 Jun;26(6):695-701 PMID: 18500337
  60. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis.
    Bioinformatics. 2005 Aug 15;21(16):3439-40 PMID: 16082012
  61. Accurate multiplex polony sequencing of an evolved bacterial genome.
    Science. 2005 Sep 9;309(5741):1728-32 PMID: 16081699
  62. Next-generation sequencing of vertebrate experimental organisms.
    Mamm Genome. 2009 Jun;20(6):327-38 PMID: 19452216
  63. Parameter estimation for robust HMM analysis of ChIP-chip data.
    BMC Bioinformatics. 2008 Aug 18;9:343 PMID: 18706106
  64. ISOLATE: a computational strategy for identifying the primary origin of cancers using high-throughput sequencing.
    Bioinformatics. 2009 Nov 1;25(21):2882-9 PMID: 19542156
  65. Rapid transcriptome characterization for a nonmodel organism using 454 pyrosequencing.
    Mol Ecol. 2008 Apr;17(7):1636-47 PMID: 18266620
  66. Real-time DNA sequencing from single polymerase molecules.
    Science. 2009 Jan 2;323(5910):133-8 PMID: 19023044
Article Info
Journal
Journal of proteomics & bioinformatics
Abbr.
J Proteomics Bioinform
ISSN
0974-276X
Published
2010-06-01
Pages
183-190
Language
English
Region
United States
NLM ID
101479045
PMCID
PMC2989618
Grants
NIEHS NIH HHS · P30 ES014443 · United States
NCI NIH HHS · R15 CA133844 · United States
NCI NIH HHS · R15 CA133844-01A2 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]