Abstract
Next-generation sequencing (NGS) technologies have generated enormous amounts of shotgun read data, and assembly of the reads can be challenging, especially for organisms without template sequences. We study the power of genome comparison based on shotgun read data without assembly using three alignment-free sequence comparison statistics, D(2), D(*)(2) and D(s)(2), both theoretically and by simulations. Theoretical formulas for the power of detecting the relationship between two sequences related through a common motif model are derived. It is shown that both D(*)(2) and D(s)(2), outperform D(2) for detecting the relationship between two sequences based on NGS data. We then study the effects of length of the tuple, read length, coverage, and sequencing error on the power of D(*)(2) and D(s)(2). Finally, variations of these statistics, d(2), d(*)(2) and d(s)(2), respectively, are used to first cluster five mammalian species with known phylogenetic relationships, and then cluster 13 tree species whose complete genome sequences are not available using NGS shotgun reads. The clustering results using d(s)(2) are consistent with biological knowledge for the 5 mammalian and 13 tree species, respectively. Thus, the statistic d(s)(2) provides a powerful alignment-free comparison tool to study the relationships among different organisms based on NGS read data without assembly.
MeSH Terms
Algorithms
Animals
Base Composition
Chickens/genetics
Cluster Analysis
Computer Simulation
Genome
High-Throughput Nucleotide Sequencing
Humans
Mice
Models, Genetic
Opossums/genetics
Phylogeny
Rabbits
Sequence Analysis, DNA/methods,statistics & numerical data
Trees/classification,genetics
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Song Kai
School of Mathematics, Peking University, Beijing, PR China.
Ren Jie
Zhai Zhiyuan
Liu Xuemei
Deng Minghua
Sun Fengzhu
References (18)
18 references, click to expand
-
28-way vertebrate alignment and conservation track in the UCSC Genome Browser.
Genome Res. 2007 Dec;17(12):1797-808
PMID: 17984227
-
Computational discovery of cis-regulatory modules in Drosophila without prior knowledge of motifs.
Genome Biol. 2008 Jan 28;9(1):R22
PMID: 18226245
-
Modeling non-uniformity in short-read rates in RNA-Seq data.
Genome Biol. 2010;11(5):R50
PMID: 20459815
-
Identifying cis-regulatory sequences by word profile similarity.
PLoS One. 2009 Sep 04;4(9):e6901
PMID: 19730735
-
Alignment-free sequence comparison (II): theoretical power of comparison statistics.
J Comput Biol. 2010 Nov;17(11):1467-90
PMID: 20973742
-
Alignment-free genome comparison with feature frequency profiles (FFP) and optimal resolutions.
Proc Natl Acad Sci U S A. 2009 Feb 24;106(8):2677-82
PMID: 19188606
-
Distributional regimes for the number of k-word matches between two random sequences.
Proc Natl Acad Sci U S A. 2002 Oct 29;99(22):13980-9
PMID: 12374863
-
New powerful statistics for alignment-free sequence comparison under a pattern transfer model.
J Theor Biol. 2011 Sep 7;284(1):106-16
PMID: 21723298
-
Modeling ChIP sequencing in silico with applications.
PLoS Comput Biol. 2008 Aug 22;4(8):e1000158
PMID: 18725927
-
Alignment-free sequence comparison (I): statistics and power.
J Comput Biol. 2009 Dec;16(12):1615-34
PMID: 20001252
-
Alignment-free detection of local similarity among viral and bacterial genomes.
Bioinformatics. 2011 Jun 1;27(11):1466-72
PMID: 21471011
-
MetaSim: a sequencing simulator for genomics and metagenomics.
PLoS One. 2008 Oct 08;3(10):e3373
PMID: 18841204
-
Assembly free comparative genomics of short-read sequence data discovers the needles in the haystack.
Mol Ecol. 2010 Mar;19 Suppl 1:147-61
PMID: 20331777
-
Alignment-free sequence comparison-a review.
Bioinformatics. 2003 Mar 1;19(4):513-23
PMID: 12611807
-
Whole-proteome phylogeny of prokaryotes by feature frequency profiles: An alignment-free method with optimal feature resolution.
Proc Natl Acad Sci U S A. 2010 Jan 5;107(1):133-8
PMID: 20018669
-
The power of detecting enriched patterns: an HMM approach.
J Comput Biol. 2010 Apr;17(4):581-92
PMID: 20426691
-
Biases in Illumina transcriptome sequencing caused by random hexamer priming.
Nucleic Acids Res. 2010 Jul;38(12):e131
PMID: 20395217
-
A measure of the similarity of sets of sequences not requiring sequence alignment.
Proc Natl Acad Sci U S A. 1986 Jul;83(14):5155-9
PMID: 3460087