Home LiteratureArticle Details
PMID: 23383994 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

Alignment-free sequence comparison based on next-generation sequencing reads.

Song K, Ren J, Zhai Z, Liu X, Deng M, Sun F

Abstract

Next-generation sequencing (NGS) technologies have generated enormous amounts of shotgun read data, and assembly of the reads can be challenging, especially for organisms without template sequences. We study the power of genome comparison based on shotgun read data without assembly using three alignment-free sequence comparison statistics, D(2), D(*)(2) and D(s)(2), both theoretically and by simulations. Theoretical formulas for the power of detecting the relationship between two sequences related through a common motif model are derived. It is shown that both D(*)(2) and D(s)(2), outperform D(2) for detecting the relationship between two sequences based on NGS data. We then study the effects of length of the tuple, read length, coverage, and sequencing error on the power of D(*)(2) and D(s)(2). Finally, variations of these statistics, d(2), d(*)(2) and d(s)(2), respectively, are used to first cluster five mammalian species with known phylogenetic relationships, and then cluster 13 tree species whose complete genome sequences are not available using NGS shotgun reads. The clustering results using d(s)(2) are consistent with biological knowledge for the 5 mammalian and 13 tree species, respectively. Thus, the statistic d(s)(2) provides a powerful alignment-free comparison tool to study the relationships among different organisms based on NGS read data without assembly.

MeSH Terms
Algorithms Animals Base Composition Chickens/genetics Cluster Analysis Computer Simulation Genome High-Throughput Nucleotide Sequencing Humans Mice Models, Genetic Opossums/genetics Phylogeny Rabbits Sequence Analysis, DNA/methods,statistics & numerical data Trees/classification,genetics
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Song Kai
School of Mathematics, Peking University, Beijing, PR China.
Ren Jie
Zhai Zhiyuan
Liu Xuemei
Deng Minghua
Sun Fengzhu
References (18)
18 references, click to expand
  1. 28-way vertebrate alignment and conservation track in the UCSC Genome Browser.
    Genome Res. 2007 Dec;17(12):1797-808 PMID: 17984227
  2. Computational discovery of cis-regulatory modules in Drosophila without prior knowledge of motifs.
    Genome Biol. 2008 Jan 28;9(1):R22 PMID: 18226245
  3. Modeling non-uniformity in short-read rates in RNA-Seq data.
    Genome Biol. 2010;11(5):R50 PMID: 20459815
  4. Identifying cis-regulatory sequences by word profile similarity.
    PLoS One. 2009 Sep 04;4(9):e6901 PMID: 19730735
  5. Alignment-free sequence comparison (II): theoretical power of comparison statistics.
    J Comput Biol. 2010 Nov;17(11):1467-90 PMID: 20973742
  6. Alignment-free genome comparison with feature frequency profiles (FFP) and optimal resolutions.
    Proc Natl Acad Sci U S A. 2009 Feb 24;106(8):2677-82 PMID: 19188606
  7. Distributional regimes for the number of k-word matches between two random sequences.
    Proc Natl Acad Sci U S A. 2002 Oct 29;99(22):13980-9 PMID: 12374863
  8. New powerful statistics for alignment-free sequence comparison under a pattern transfer model.
    J Theor Biol. 2011 Sep 7;284(1):106-16 PMID: 21723298
  9. Modeling ChIP sequencing in silico with applications.
    PLoS Comput Biol. 2008 Aug 22;4(8):e1000158 PMID: 18725927
  10. Alignment-free sequence comparison (I): statistics and power.
    J Comput Biol. 2009 Dec;16(12):1615-34 PMID: 20001252
  11. Alignment-free detection of local similarity among viral and bacterial genomes.
    Bioinformatics. 2011 Jun 1;27(11):1466-72 PMID: 21471011
  12. MetaSim: a sequencing simulator for genomics and metagenomics.
    PLoS One. 2008 Oct 08;3(10):e3373 PMID: 18841204
  13. Assembly free comparative genomics of short-read sequence data discovers the needles in the haystack.
    Mol Ecol. 2010 Mar;19 Suppl 1:147-61 PMID: 20331777
  14. Alignment-free sequence comparison-a review.
    Bioinformatics. 2003 Mar 1;19(4):513-23 PMID: 12611807
  15. Whole-proteome phylogeny of prokaryotes by feature frequency profiles: An alignment-free method with optimal feature resolution.
    Proc Natl Acad Sci U S A. 2010 Jan 5;107(1):133-8 PMID: 20018669
  16. The power of detecting enriched patterns: an HMM approach.
    J Comput Biol. 2010 Apr;17(4):581-92 PMID: 20426691
  17. Biases in Illumina transcriptome sequencing caused by random hexamer priming.
    Nucleic Acids Res. 2010 Jul;38(12):e131 PMID: 20395217
  18. A measure of the similarity of sets of sequences not requiring sequence alignment.
    Proc Natl Acad Sci U S A. 1986 Jul;83(14):5155-9 PMID: 3460087
Article Info
Journal
Journal of computational biology : a journal of computational molecular cell biology
Abbr.
J Comput Biol
ISSN
1557-8666
Published
2013-02-00
Pages
64-79
Language
English
Region
United States
NLM ID
9433358
PMCID
PMC3581251
Subset
IM
Grants
NHGRI NIH HHS · P50 HG 002790 · United States
NHGRI NIH HHS · R21HG006199 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]