Home LiteratureArticle Details
PMID: 20109217 Published · epublish English Evaluation Study Journal Article Research Support, Non-U.S. Gov't Validation Study

Validation of differential gene expression algorithms: application comparing fold-change estimation to hypothesis testing.

BMC bioinformatics ·Vol. 11 ·2010-01-28 ·Pages 63

Yanofsky CM, Bickel DR

Abstract

Sustained research on the problem of determining which genes are differentially expressed on the basis of microarray data has yielded a plethora of statistical algorithms, each justified by theory, simulation, or ad hoc validation and yet differing in practical results from equally justified algorithms. Recently, a concordance method that measures agreement among gene lists have been introduced to assess various aspects of differential gene expression detection. This method has the advantage of basing its assessment solely on the results of real data analyses, but as it requires examining gene lists of given sizes, it may be unstable. Two methodologies for assessing predictive error are described: a cross-validation method and a posterior predictive method. As a nonparametric method of estimating prediction error from observed expression levels, cross validation provides an empirical approach to assessing algorithms for detecting differential gene expression that is fully justified for large numbers of biological replicates. Because it leverages the knowledge that only a small portion of genes are differentially expressed, the posterior predictive method is expected to provide more reliable estimates of algorithm performance, allaying concerns about limited biological replication. In practice, the posterior predictive method can assess when its approximations are valid and when they are inaccurate. Under conditions in which its approximations are valid, it corroborates the results of cross validation. Both comparison methodologies are applicable to both single-channel and dual-channel microarrays. For the data sets considered, estimating prediction error by cross validation demonstrates that empirical Bayes methods based on hierarchical models tend to outperform algorithms based on selecting genes by their fold changes or by non-hierarchical model-selection criteria. (The latter two approaches have comparable performance.) The posterior predictive assessment corroborates these findings. Algorithms for detecting differential gene expression may be compared by estimating each algorithm's error in predicting expression ratios, whether such ratios are defined across microarray channels or between two independent groups.According to two distinct estimators of prediction error, algorithms using hierarchical models outperform the other algorithms of the study. The fact that fold-change shrinkage performed as well as conventional model selection criteria calls for investigating algorithms that combine the strengths of significance testing and fold-change estimation.

MeSH Terms
Algorithms Data Interpretation, Statistical Gene Expression Profiling/methods Oligonucleotide Array Sequence Analysis/methods
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Yanofsky Corey M
Ottawa Institute of Systems Biology, Department of Biochemistry, Microbiology, and Immunology, University of Ottawa, Ottawa, Ontario, Canada.
Bickel David R
References (30)
30 references, click to expand
  1. Multidimensional local false discovery rate for microarray studies.
    Bioinformatics. 2006 Mar 1;22(5):556-65 PMID: 16368770
  2. Tail posterior probability for inference in pairwise and multiclass gene expression data.
    Biometrics. 2007 Dec;63(4):1117-25 PMID: 18078482
  3. Error-rate and decision-theoretic methods of multiple testing: which genes have high objective probabilities of differential expression?
    Stat Appl Genet Mol Biol. 2004;3:Article8 PMID: 16646824
  4. Degrees of differential gene expression: detecting biologically significant expression differences and estimating their magnitudes.
    Bioinformatics. 2004 Mar 22;20(5):682-8 PMID: 15033875
  5. Confirming microarray data--is it really necessary?
    Genomics. 2004 Apr;83(4):541-9 PMID: 15028276
  6. Consolidated strategy for the analysis of microarray spike-in data.
    Nucleic Acids Res. 2008 Oct;36(17):e108 PMID: 18676452
  7. A comparison of methods to control type I errors in microarray studies.
    Stat Appl Genet Mol Biol. 2007;6:Article28 PMID: 18052911
  8. Estimating the occurrence of false positives and false negatives in microarray studies by approximating and partitioning the empirical distribution of p-values.
    Bioinformatics. 2003 Jul 1;19(10):1236-42 PMID: 12835267
  9. Comparison of small n statistical tests of differential expression applied to microarrays.
    BMC Bioinformatics. 2009 Feb 03;10:45 PMID: 19192265
  10. Microarray data analysis: from disarray to consolidation and consensus.
    Nat Rev Genet. 2006 Jan;7(1):55-65 PMID: 16369572
  11. Empirical evaluation of data transformations and ranking statistics for microarray analysis.
    Nucleic Acids Res. 2004 Oct 12;32(18):5471-9 PMID: 15479783
  12. Estimating the false discovery rate using nonparametric deconvolution.
    Biometrics. 2007 Sep;63(3):806-15 PMID: 17825012
  13. Determination of the differentially expressed genes in microarray experiments using local FDR.
    BMC Bioinformatics. 2004 Sep 06;5:125 PMID: 15350197
  14. Rat toxicogenomic study reveals analytical consistency across microarray platforms.
    Nat Biotechnol. 2006 Sep;24(9):1162-9 PMID: 17061323
  15. Reproducibility of microarray data: a further analysis of microarray quality control (MAQC) data.
    BMC Bioinformatics. 2007 Oct 25;8:412 PMID: 17961233
  16. twilight; a Bioconductor package for estimating the local false discovery rate.
    Bioinformatics. 2005 Jun 15;21(12):2921-2 PMID: 15817688
  17. Selection of differentially expressed genes in microarray data analysis.
    Pharmacogenomics J. 2007 Jun;7(3):212-20 PMID: 16940966
  18. Bayesian modeling of differential gene expression.
    Biometrics. 2006 Mar;62(1):1-9 PMID: 16542224
  19. The balance of reproducibility, sensitivity, and specificity of lists of differentially expressed genes in microarray studies.
    BMC Bioinformatics. 2008 Aug 12;9 Suppl 9:S10 PMID: 18793455
  20. Transcriptome and selected metabolite analyses reveal multiple points of ethylene control during tomato fruit development.
    Plant Cell. 2005 Nov;17(11):2954-65 PMID: 16243903
  21. Mixture models for detecting differentially expressed genes in microarrays.
    Int J Neural Syst. 2006 Oct;16(5):353-62 PMID: 17117496
  22. Bioconductor: open software development for computational biology and bioinformatics.
    Genome Biol. 2004;5(10):R80 PMID: 15461798
  23. Selecting differentially expressed genes from microarray experiments.
    Biometrics. 2003 Mar;59(1):133-42 PMID: 12762450
  24. A stochastic downhill search algorithm for estimating the local false discovery rate.
    IEEE/ACM Trans Comput Biol Bioinform. 2004 Jul-Sep;1(3):98-108 PMID: 17048385
  25. A mixture model for estimating the local false discovery rate in DNA microarray analysis.
    Bioinformatics. 2004 Nov 1;20(16):2694-701 PMID: 15145810
  26. Significance testing for small microarray experiments.
    Stat Med. 2005 Aug 15;24(15):2281-98 PMID: 15889452
  27. Correcting the estimated level of differential expression for gene selection bias: application to a microarray study.
    Stat Appl Genet Mol Biol. 2008;7(1):Article10 PMID: 18384263
  28. Testing significance relative to a fold-change threshold is a TREAT.
    Bioinformatics. 2009 Mar 15;25(6):765-71 PMID: 19176553
  29. Comparison and evaluation of methods for generating differentially expressed gene lists from microarray data.
    BMC Bioinformatics. 2006 Jul 26;7:359 PMID: 16872483
  30. Linear models and empirical bayes methods for assessing differential expression in microarray experiments.
    Stat Appl Genet Mol Biol. 2004;3:Article3 PMID: 16646809
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2010-01-28
Epub
2010-00-28
Pages
63
Language
English
Region
England
NLM ID
100965194
PMCID
PMC3224549
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]