Home LiteratureArticle Details
PMID: 19750212 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Fast mapping of short sequences with mismatches, insertions and deletions using index structures.

PLoS computational biology ·Vol. 5 ·No. 9 ·2009-09-00 ·Pages e1000502

Hoffmann S, Otto C, Kurtz S, Sharma CM, Khaitovich P, Vogel J, Stadler PF, Hackermüller J

Abstract

With few exceptions, current methods for short read mapping make use of simple seed heuristics to speed up the search. Most of the underlying matching models neglect the necessity to allow not only mismatches, but also insertions and deletions. Current evaluations indicate, however, that very different error models apply to the novel high-throughput sequencing methods. While the most frequent error-type in Illumina reads are mismatches, reads produced by 454's GS FLX predominantly contain insertions and deletions (indels). Even though 454 sequencers are able to produce longer reads, the method is frequently applied to small RNA (miRNA and siRNA) sequencing. Fast and accurate matching in particular of short reads with diverse errors is therefore a pressing practical problem. We introduce a matching model for short reads that can, besides mismatches, also cope with indels. It addresses different error models. For example, it can handle the problem of leading and trailing contaminations caused by primers and poly-A tails in transcriptomics or the length-dependent increase of error rates. In these contexts, it thus simplifies the tedious and error-prone trimming step. For efficient searches, our method utilizes index structures in the form of enhanced suffix arrays. In a comparison with current methods for short read mapping, the presented approach shows significantly increased performance not only for 454 reads, but also for Illumina reads. Our approach is implemented in the software segemehl available at http://www.bioinf.uni-leipzig.de/Software/segemehl/.

MeSH Terms
Algorithms Base Sequence Computational Biology/methods DNA Mutational Analysis/methods Mutation Sequence Alignment
Authors & Affiliations
8 authors, click to expand affiliations / ORCID
Hoffmann Steve
Bioinformatics Group, Department of Computer Science, University of Leipzig, Leipzig, Germany.
Otto Christian
Kurtz Stefan
Sharma Cynthia M
Khaitovich Philipp
Vogel Jörg
Stadler Peter F
Hackermüller Jörg
References (12)
12 references, click to expand
  1. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  2. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  3. PatMaN: rapid alignment of short sequences to large databases.
    Bioinformatics. 2008 Jul 1;24(13):1530-1 PMID: 18467344
  4. SHRiMP: accurate mapping of short color-space reads.
    PLoS Comput Biol. 2009 May;5(5):e1000386 PMID: 19461883
  5. SOAP: short oligonucleotide alignment program.
    Bioinformatics. 2008 Mar 1;24(5):713-4 PMID: 18227114
  6. Accuracy and quality of massively parallel DNA pyrosequencing.
    Genome Biol. 2007;8(7):R143 PMID: 17659080
  7. Mapping short DNA sequencing reads and calling variants using mapping quality scores.
    Genome Res. 2008 Nov;18(11):1851-8 PMID: 18714091
  8. The development and impact of 454 sequencing.
    Nat Biotechnol. 2008 Oct;26(10):1117-24 PMID: 18846085
  9. ZOOM! Zillions of oligos mapped.
    Bioinformatics. 2008 Nov 1;24(21):2431-7 PMID: 18684737
  10. Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes.
    Proc Natl Acad Sci U S A. 1990 Mar;87(6):2264-8 PMID: 2315319
  11. Solexa Ltd.
    Pharmacogenomics. 2004 Jun;5(4):433-8 PMID: 15165179
  12. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
Article Info
Journal
PLoS computational biology
Abbr.
PLoS Comput Biol
ISSN
1553-7358
Published
2009-09-00
Epub
2009-00-11
Pages
e1000502
Language
English
Region
United States
NLM ID
101238922
PMCID
PMC2730575
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]