Abstract
Sustained research on the problem of determining which genes are differentially expressed on the basis of microarray data has yielded a plethora of statistical algorithms, each justified by theory, simulation, or ad hoc validation and yet differing in practical results from equally justified algorithms. Recently, a concordance method that measures agreement among gene lists have been introduced to assess various aspects of differential gene expression detection. This method has the advantage of basing its assessment solely on the results of real data analyses, but as it requires examining gene lists of given sizes, it may be unstable. Two methodologies for assessing predictive error are described: a cross-validation method and a posterior predictive method. As a nonparametric method of estimating prediction error from observed expression levels, cross validation provides an empirical approach to assessing algorithms for detecting differential gene expression that is fully justified for large numbers of biological replicates. Because it leverages the knowledge that only a small portion of genes are differentially expressed, the posterior predictive method is expected to provide more reliable estimates of algorithm performance, allaying concerns about limited biological replication. In practice, the posterior predictive method can assess when its approximations are valid and when they are inaccurate. Under conditions in which its approximations are valid, it corroborates the results of cross validation. Both comparison methodologies are applicable to both single-channel and dual-channel microarrays. For the data sets considered, estimating prediction error by cross validation demonstrates that empirical Bayes methods based on hierarchical models tend to outperform algorithms based on selecting genes by their fold changes or by non-hierarchical model-selection criteria. (The latter two approaches have comparable performance.) The posterior predictive assessment corroborates these findings. Algorithms for detecting differential gene expression may be compared by estimating each algorithm's error in predicting expression ratios, whether such ratios are defined across microarray channels or between two independent groups.According to two distinct estimators of prediction error, algorithms using hierarchical models outperform the other algorithms of the study. The fact that fold-change shrinkage performed as well as conventional model selection criteria calls for investigating algorithms that combine the strengths of significance testing and fold-change estimation.
MeSH Terms
Algorithms
Data Interpretation, Statistical
Gene Expression Profiling/methods
Oligonucleotide Array Sequence Analysis/methods
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Yanofsky Corey M
Ottawa Institute of Systems Biology, Department of Biochemistry, Microbiology, and Immunology, University of Ottawa, Ottawa, Ontario, Canada.
Bickel David R
References (30)
30 references, click to expand
-
Multidimensional local false discovery rate for microarray studies.
Bioinformatics. 2006 Mar 1;22(5):556-65
PMID: 16368770
-
Tail posterior probability for inference in pairwise and multiclass gene expression data.
Biometrics. 2007 Dec;63(4):1117-25
PMID: 18078482
-
Error-rate and decision-theoretic methods of multiple testing: which genes have high objective probabilities of differential expression?
Stat Appl Genet Mol Biol. 2004;3:Article8
PMID: 16646824
-
Degrees of differential gene expression: detecting biologically significant expression differences and estimating their magnitudes.
Bioinformatics. 2004 Mar 22;20(5):682-8
PMID: 15033875
-
Confirming microarray data--is it really necessary?
Genomics. 2004 Apr;83(4):541-9
PMID: 15028276
-
Consolidated strategy for the analysis of microarray spike-in data.
Nucleic Acids Res. 2008 Oct;36(17):e108
PMID: 18676452
-
A comparison of methods to control type I errors in microarray studies.
Stat Appl Genet Mol Biol. 2007;6:Article28
PMID: 18052911
-
Estimating the occurrence of false positives and false negatives in microarray studies by approximating and partitioning the empirical distribution of p-values.
Bioinformatics. 2003 Jul 1;19(10):1236-42
PMID: 12835267
-
Comparison of small n statistical tests of differential expression applied to microarrays.
BMC Bioinformatics. 2009 Feb 03;10:45
PMID: 19192265
-
Microarray data analysis: from disarray to consolidation and consensus.
Nat Rev Genet. 2006 Jan;7(1):55-65
PMID: 16369572
-
Empirical evaluation of data transformations and ranking statistics for microarray analysis.
Nucleic Acids Res. 2004 Oct 12;32(18):5471-9
PMID: 15479783
-
Estimating the false discovery rate using nonparametric deconvolution.
Biometrics. 2007 Sep;63(3):806-15
PMID: 17825012
-
Determination of the differentially expressed genes in microarray experiments using local FDR.
BMC Bioinformatics. 2004 Sep 06;5:125
PMID: 15350197
-
Rat toxicogenomic study reveals analytical consistency across microarray platforms.
Nat Biotechnol. 2006 Sep;24(9):1162-9
PMID: 17061323
-
Reproducibility of microarray data: a further analysis of microarray quality control (MAQC) data.
BMC Bioinformatics. 2007 Oct 25;8:412
PMID: 17961233
-
twilight; a Bioconductor package for estimating the local false discovery rate.
Bioinformatics. 2005 Jun 15;21(12):2921-2
PMID: 15817688
-
Selection of differentially expressed genes in microarray data analysis.
Pharmacogenomics J. 2007 Jun;7(3):212-20
PMID: 16940966
-
Bayesian modeling of differential gene expression.
Biometrics. 2006 Mar;62(1):1-9
PMID: 16542224
-
The balance of reproducibility, sensitivity, and specificity of lists of differentially expressed genes in microarray studies.
BMC Bioinformatics. 2008 Aug 12;9 Suppl 9:S10
PMID: 18793455
-
Transcriptome and selected metabolite analyses reveal multiple points of ethylene control during tomato fruit development.
Plant Cell. 2005 Nov;17(11):2954-65
PMID: 16243903
-
Mixture models for detecting differentially expressed genes in microarrays.
Int J Neural Syst. 2006 Oct;16(5):353-62
PMID: 17117496
-
Bioconductor: open software development for computational biology and bioinformatics.
Genome Biol. 2004;5(10):R80
PMID: 15461798
-
Selecting differentially expressed genes from microarray experiments.
Biometrics. 2003 Mar;59(1):133-42
PMID: 12762450
-
A stochastic downhill search algorithm for estimating the local false discovery rate.
IEEE/ACM Trans Comput Biol Bioinform. 2004 Jul-Sep;1(3):98-108
PMID: 17048385
-
A mixture model for estimating the local false discovery rate in DNA microarray analysis.
Bioinformatics. 2004 Nov 1;20(16):2694-701
PMID: 15145810
-
Significance testing for small microarray experiments.
Stat Med. 2005 Aug 15;24(15):2281-98
PMID: 15889452
-
Correcting the estimated level of differential expression for gene selection bias: application to a microarray study.
Stat Appl Genet Mol Biol. 2008;7(1):Article10
PMID: 18384263
-
Testing significance relative to a fold-change threshold is a TREAT.
Bioinformatics. 2009 Mar 15;25(6):765-71
PMID: 19176553
-
Comparison and evaluation of methods for generating differentially expressed gene lists from microarray data.
BMC Bioinformatics. 2006 Jul 26;7:359
PMID: 16872483
-
Linear models and empirical bayes methods for assessing differential expression in microarray experiments.
Stat Appl Genet Mol Biol. 2004;3:Article3
PMID: 16646809