Home LiteratureArticle Details
PMID: 20724458 Published · ppublish English Journal Article Review

De novo assembly of short sequence reads.

Briefings in bioinformatics ·Vol. 11 ·No. 5 ·2010-09-00 ·Pages 457-72

Paszkiewicz K, Studholme DJ

Abstract

A new generation of sequencing technologies is revolutionizing molecular biology. Illumina's Solexa and Applied Biosystems' SOLiD generate gigabases of nucleotide sequence per week. However, a perceived limitation of these ultra-high-throughput technologies is their short read-lengths. De novo assembly of sequence reads generated by classical Sanger capillary sequencing is a mature field of research. Unfortunately, the existing sequence assembly programs were not effective for short sequence reads generated by Illumina and SOLiD platforms. Early studies suggested that, in principle, sequence reads as short as 20-30 nucleotides could be used to generate useful assemblies of both prokaryotic and eukaryotic genome sequences, albeit containing many gaps. The early feasibility studies and proofs of principle inspired several bioinformatics research groups to implement new algorithms as freely available software tools specifically aimed at assembling reads of 30-50 nucleotides in length. This has led to the generation of several draft genome sequences based exclusively on short sequence Illumina sequence reads, recently culminating in the assembly of the 2.25-Gb genome of the giant panda from Illumina sequence reads with an average length of just 52 nucleotides. As well as reviewing recent developments in the field, we discuss some practical aspects such as data filtering and submission of assembly data to public repositories.

MeSH Terms
Algorithms Animals Base Sequence Computational Biology/methods Databases, Genetic Genome Humans Molecular Sequence Data Sequence Analysis, DNA/methods
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Paszkiewicz Konrad
Imperial College, London.
Studholme David J
Article Info
Journal
Briefings in bioinformatics
Abbr.
Brief Bioinform
ISSN
1477-4054
Published
2010-09-00
Epub
2010-00-19
Pages
457-72
Language
English
Region
England
NLM ID
100912837
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]