Home LiteratureArticle Details
PMID: 15262778 Published · ppublish English Evaluation Study Journal Article Research Support, Non-U.S. Gov't

Statistical modeling of sequencing errors in SAGE libraries.

Bioinformatics (Oxford, England) ·Vol. 20 Suppl 1 ·2004-08-04 ·Pages i31-9

Beissbarth T, Hyde L, Smyth GK, Job C, Boon WM, Tan SS, Scott HS, Speed TP

Abstract

Sequencing errors may bias the gene expression measurements made by Serial Analysis of Gene Expression (SAGE). They may introduce non-existent tags at low abundance and decrease the real abundance of other tags. These effects are increased in the longer tags generated in LongSAGE libraries. Current sequencing technology generates quite accurate estimates of sequencing error rates. Here we make use of the sequence neighborhood of SAGE tags and error estimates from the base-calling software to correct for such errors. We introduce a statistical model for the propagation of sequencing errors in SAGE and suggest an Expectation-Maximization (EM) algorithm to correct for them given observed sequences in a library and base-calling error estimates. We tested our method using simulated and experimental SAGE libraries. When comparing SAGE libraries, we found that sequencing errors can introduce considerable bias. High abundance tags may be falsely called as significantly differentially expressed, especially when comparing libraries with different levels of sequencing errors and/or of different size. Truly, differentially expressed tags have decreased significance as 'true'-tag counts are generally underestimated. This may alter if tags near the threshold of differential expression are called significant. Moreover, the number of different transcripts present in a library is overestimated as false tags are introduced at low abundance. Our correction method adjusts the tag counts to be closer to the true counts and is able to partly correct for biases introduced by sequencing errors. An implementation using R is distributed as an R package. An online version is available at http://tagcalling.mbgproject.org

MeSH Terms
Algorithms Base Sequence Computer Simulation Data Interpretation, Statistical Expressed Sequence Tags Gene Expression Profiling/methods Gene Library Models, Genetic Models, Statistical Molecular Sequence Data Reproducibility of Results Sensitivity and Specificity Sequence Analysis, DNA/methods
Authors & Affiliations
8 authors, click to expand affiliations / ORCID
Beissbarth Tim
Walter and Eliza Hall Institute of Medical Research, Genetics and Bioinformatics, Parkville, Vic, Australia. [email protected]
Hyde Lavinia
Smyth Gordon K
Job Chris
Boon Wee-Ming
Tan Seong-Seng
Scott Hamish S
Speed Terence P
Article Info
Journal
Bioinformatics (Oxford, England)
Abbr.
Bioinformatics
ISSN
1367-4811
Published
2004-08-04
Pages
i31-9
Language
English
Region
England
NLM ID
9808944
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]