Home LiteratureArticle Details
PMID: 18851737 Published · epublish English Journal Article Research Support, Non-U.S. Gov't

Probabilistic base calling of Solexa sequencing data.

BMC bioinformatics ·Vol. 9 ·2008-10-13 ·Pages 431

Rougemont J, Amzallag A, Iseli C, Farinelli L, Xenarios I, Naef F

Abstract

Solexa/Illumina short-read ultra-high throughput DNA sequencing technology produces millions of short tags (up to 36 bases) by parallel sequencing-by-synthesis of DNA colonies. The processing and statistical analysis of such high-throughput data poses new challenges; currently a fair proportion of the tags are routinely discarded due to an inability to match them to a reference sequence, thereby reducing the effective throughput of the technology. We propose a novel base calling algorithm using model-based clustering and probability theory to identify ambiguous bases and code them with IUPAC symbols. We also select optimal sub-tags using a score based on information content to remove uncertain bases towards the ends of the reads. We show that the method improves genome coverage and number of usable tags as compared with Solexa's data processing pipeline by an average of 15%. An R package is provided which allows fast and accurate base calling of Solexa's fluorescence intensity files and the production of informative diagnostic plots.

MeSH Terms
Bacteriophage phi X 174/genetics Base Sequence/genetics Chromosome Mapping/methods Cluster Analysis DNA, Viral/analysis Expressed Sequence Tags Pattern Recognition, Automated/methods Quality Control Sequence Analysis, DNA/methods Software Spectrometry, Fluorescence/methods
Chemicals
DNA, Viral
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Rougemont Jacques
School of Life Sciences, Ecole Polytechnique Fédérale de Lausanne (EPFL), 1015 Lausanne, Switzerland. [email protected]
Amzallag Arnaud
Iseli Christian
Farinelli Laurent
Xenarios Ioannis
Naef Felix
References (22)
22 references, click to expand
  1. Optimal alignments in linear space.
    Comput Appl Biosci. 1988 Mar;4(1):11-7 PMID: 3382986
  2. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
  3. Whole-genome patterns of common DNA variation in three human populations.
    Science. 2005 Feb 18;307(5712):1072-9 PMID: 15718463
  4. Genome sequencing in microfabricated high-density picolitre reactors.
    Nature. 2005 Sep 15;437(7057):376-80 PMID: 16056220
  5. Base-stacking and base-pairing contributions into thermal stability of the DNA double helix.
    Nucleic Acids Res. 2006;34(2):564-74 PMID: 16449200
  6. Whole-genome re-sequencing.
    Curr Opin Genet Dev. 2006 Dec;16(6):545-52 PMID: 17055251
  7. High-resolution profiling of histone methylations in the human genome.
    Cell. 2007 May 18;129(4):823-37 PMID: 17512414
  8. Indexing strategies for rapid searches of short words in genome sequences.
    PLoS One. 2007;2(6):e579 PMID: 17593978
  9. Optimized design and assessment of whole genome tiling arrays.
    Bioinformatics. 2007 Jul 1;23(13):i195-204 PMID: 17646297
  10. Genome-wide maps of chromatin state in pluripotent and lineage-committed cells.
    Nature. 2007 Aug 2;448(7153):553-60 PMID: 17603471
  11. Paired-end mapping reveals extensive structural variation in the human genome.
    Science. 2007 Oct 19;318(5849):420-6 PMID: 17901297
  12. Identification of microRNAs and other small regulatory RNAs using cDNA library sequencing.
    Methods. 2008 Jan;44(1):3-12 PMID: 18158127
  13. Bioinformatics challenges of new sequencing technology.
    Trends Genet. 2008 Mar;24(3):142-9 PMID: 18262676
  14. Shotgun bisulphite sequencing of the Arabidopsis genome reveals DNA methylation patterning.
    Nature. 2008 Mar 13;452(7184):215-9 PMID: 18278030
  15. Using quality scores and longer reads improves accuracy of Solexa read mapping.
    BMC Bioinformatics. 2008;9:128 PMID: 18307793
  16. Rapid transcriptome characterization for a nonmodel organism using 454 pyrosequencing.
    Mol Ecol. 2008 Apr;17(7):1636-47 PMID: 18266620
  17. Discovering microRNAs from deep sequencing data using miRDeep.
    Nat Biotechnol. 2008 Apr;26(4):407-15 PMID: 18392026
  18. De novo bacterial genome sequencing: millions of very short reads assembled on a desktop computer.
    Genome Res. 2008 May;18(5):802-9 PMID: 18332092
  19. Mapping translocation breakpoints by next-generation sequencing.
    Genome Res. 2008 Jul;18(7):1143-9 PMID: 18326688
  20. TileQC: a system for tile-based quality control of Solexa data.
    BMC Bioinformatics. 2008;9:250 PMID: 18507856
  21. Alta-Cyclic: a self-optimizing base caller for next-generation sequencing.
    Nat Methods. 2008 Aug;5(8):679-82 PMID: 18604217
  22. Base-calling of automated sequencer traces using phred. II. Error probabilities.
    Genome Res. 1998 Mar;8(3):186-94 PMID: 9521922
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2008-10-13
Epub
2008-00-13
Pages
431
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2575221
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]