Abstract
Eugene Myers in his string graph paper suggested that in a string graph or equivalently a unitig graph, any path spells a valid assembly. As a string/unitig graph also encodes every valid assembly of reads, such a graph, provided that it can be constructed correctly, is in fact a lossless representation of reads. In principle, every analysis based on whole-genome shotgun sequencing (WGS) data, such as SNP and insertion/deletion (INDEL) calling, can also be achieved with unitigs. To explore the feasibility of using de novo assembly in the context of resequencing, we developed a de novo assembler, fermi, that assembles Illumina short reads into unitigs while preserving most of information of the input reads. SNPs and INDELs can be called by mapping the unitigs against a reference genome. By applying the method on 35-fold human resequencing data, we showed that in comparison to the standard pipeline, our approach yields similar accuracy for SNP calling and better results for INDEL calling. It has higher sensitivity than other de novo assembly based methods for variant calling. Our work suggests that variant calling with de novo assembly can be a beneficial complement to the standard variant calling pipeline for whole-genome resequencing. In the methodological aspects, we propose FMD-index for forward-backward extension of DNA sequences, a fast algorithm for finding all super-maximal exact matches and one-pass construction of unitigs from an FMD-index. http://github.com/lh3/fermi
MeSH Terms
Algorithms
Computational Biology/methods
Humans
INDEL Mutation
Polymorphism, Single Nucleotide
Sequence Analysis, DNA/methods
Authors & Affiliations
1 authors, click to expand affiliations / ORCID
Li Heng
Medical Population Genetics Program, Broad Institute, 7 Cambridge Center, MA 02142, USA.
[email protected]
References (31)
31 references, click to expand
-
Fast and accurate short read alignment with Burrows-Wheeler transform.
Bioinformatics. 2009 Jul 15;25(14):1754-60
PMID: 19451168
-
Efficient construction of an assembly string graph using the FM-index.
Bioinformatics. 2010 Jun 15;26(12):i367-73
PMID: 20529929
-
Computational techniques for human genome resequencing using mated gapped reads.
J Comput Biol. 2012 Mar;19(3):279-92
PMID: 22175250
-
SNP-o-matic.
Bioinformatics. 2009 Sep 15;25(18):2434-5
PMID: 19574284
-
The diploid genome sequence of an individual human.
PLoS Biol. 2007 Sep 4;5(10):e254
PMID: 17803354
-
Computer programs for the assembly of DNA sequences.
Nucleic Acids Res. 1979 Sep 25;7(2):529-45
PMID: 493154
-
An Eulerian path approach to DNA fragment assembly.
Proc Natl Acad Sci U S A. 2001 Aug 14;98(17):9748-53
PMID: 11504945
-
High-quality draft assemblies of mammalian genomes from massively parallel sequence data.
Proc Natl Acad Sci U S A. 2011 Jan 25;108(4):1513-8
PMID: 21187386
-
Human genome sequencing using unchained base reads on self-assembling DNA nanoarrays.
Science. 2010 Jan 1;327(5961):78-81
PMID: 19892942
-
Performance comparison of whole-genome sequencing platforms.
Nat Biotechnol. 2011 Dec 18;30(1):78-82
PMID: 22178993
-
De novo assembly and genotyping of variants using colored de Bruijn graphs.
Nat Genet. 2012 Jan 08;44(2):226-32
PMID: 22231483
-
ABySS: a parallel assembler for short read sequence data.
Genome Res. 2009 Jun;19(6):1117-23
PMID: 19251739
-
Improving SNP discovery by base alignment quality.
Bioinformatics. 2011 Apr 15;27(8):1157-8
PMID: 21320865
-
A whole-genome assembly of Drosophila.
Science. 2000 Mar 24;287(5461):2196-204
PMID: 10731133
-
The fragment assembly string graph.
Bioinformatics. 2005 Sep 1;21 Suppl 2:ii79-85
PMID: 16204131
-
A map of human genome variation from population-scale sequencing.
Nature. 2010 Oct 28;467(7319):1061-73
PMID: 20981092
-
Dindel: accurate indel calls from short-read data.
Genome Res. 2011 Jun;21(6):961-73
PMID: 20980555
-
Efficient de novo assembly of large genomes using compressed data structures.
Genome Res. 2012 Mar;22(3):549-56
PMID: 22156294
-
Fast and accurate long-read alignment with Burrows-Wheeler transform.
Bioinformatics. 2010 Mar 1;26(5):589-95
PMID: 20080505
-
A strategy of DNA sequencing employing computer programs.
Nucleic Acids Res. 1979 Jun 11;6(7):2601-10
PMID: 461197
-
Improved variant discovery through local re-alignment of short-read next-generation sequencing data using SRMA.
Genome Biol. 2010;11(10):R99
PMID: 20932289
-
HiTEC: accurate error correction in high-throughput sequencing data.
Bioinformatics. 2011 Feb 1;27(3):295-302
PMID: 21115437
-
Sequencing of natural strains of Arabidopsis thaliana with short reads.
Genome Res. 2008 Dec;18(12):2024-33
PMID: 18818371
-
A new algorithm for DNA sequence assembly.
J Comput Biol. 1995 Summer;2(2):291-306
PMID: 7497130
-
SEQAID: a DNA sequence assembling program based on a mathematical model.
Nucleic Acids Res. 1984 Jan 11;12(1 Pt 1):307-21
PMID: 6320092
-
A framework for variation discovery and genotyping using next-generation DNA sequencing data.
Nat Genet. 2011 May;43(5):491-8
PMID: 21478889
-
Toward simplifying and accurately formulating fragment assembly.
J Comput Biol. 1995 Summer;2(2):275-90
PMID: 7497129
-
De novo assembly of human genomes with massively parallel short read sequencing.
Genome Res. 2010 Feb;20(2):265-72
PMID: 20019144
-
Natural genetic variation caused by small insertions and deletions in the human genome.
Genome Res. 2011 Jun;21(6):830-9
PMID: 21460062
-
De novo fragment assembly with short mate-paired reads: Does the read length matter?
Genome Res. 2009 Feb;19(2):336-46
PMID: 19056694
-
Pebble and rock band: heuristic resolution of repeats and scaffolding in the velvet short-read de novo assembler.
PLoS One. 2009 Dec 22;4(12):e8407
PMID: 20027311