Abstract
A major goal of large-scale genomics projects is to enable the use of data from high-throughput experimental methods to predict complex phenotypes such as disease susceptibility. The DREAM5 Systems Genetics B Challenge solicited algorithms to predict soybean plant resistance to the pathogen Phytophthora sojae from training sets including phenotype, genotype, and gene expression data. The challenge test set was divided into three subcategories, one requiring prediction based on only genotype data, another on only gene expression data, and the third on both genotype and gene expression data. Here we present our approach, primarily using regularized regression, which received the best-performer award for subchallenge B2 (gene expression only). We found that despite the availability of 941 genotype markers and 28,395 gene expression features, optimal models determined by cross-validation experiments typically used fewer than ten predictors, underscoring the importance of strong regularization in noisy datasets with far more features than samples. We also present substantial analysis of the training and test setup of the challenge, identifying high variance in performance on the gold standard test sets.
MeSH Terms
Genotype
Phenotype
Phytophthora/physiology
Soybeans/genetics,microbiology
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Loh Po-Ru
Department of Mathematics and Computer Science and Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, Massachusetts, USA.
Tucker George
Berger Bonnie
References (18)
18 references, click to expand
-
Genetics of gene expression and its effect on disease.
Nature. 2008 Mar 27;452(7186):423-8
PMID: 18344981
-
Regularization Paths for Generalized Linear Models via Coordinate Descent.
J Stat Softw. 2010;33(1):1-22
PMID: 20808728
-
Gene expression profiling predicts clinical outcome of breast cancer.
Nature. 2002 Jan 31;415(6871):530-6
PMID: 11823860
-
Discovery of meaningful associations in genomic data using partial correlation coefficients.
Bioinformatics. 2004 Dec 12;20(18):3565-74
PMID: 15284096
-
Harnessing gene expression to identify the genetic basis of drug resistance.
Mol Syst Biol. 2009;5:310
PMID: 19888205
-
Infection and genotype remodel the entire soybean transcriptome.
BMC Genomics. 2009 Jan 26;10:49
PMID: 19171053
-
Reverse engineering the genotype-phenotype map with natural genetic variation.
Nature. 2008 Dec 11;456(7223):738-44
PMID: 19079051
-
Systems biology, proteomics, and the future of health care: toward predictive, preventative, and personalized medicine.
J Proteome Res. 2004 Mar-Apr;3(2):179-96
PMID: 15113093
-
Pharmacogenetics and drug development: the path to safer and more effective drugs.
Nat Rev Genet. 2004 Sep;5(9):645-56
PMID: 15372086
-
Molecular classification of cancer: class discovery and class prediction by gene expression monitoring.
Science. 1999 Oct 15;286(5439):531-7
PMID: 10521349
-
Genetic dissection of transcriptional regulation in budding yeast.
Science. 2002 Apr 26;296(5568):752-5
PMID: 11923494
-
Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling.
Nature. 2000 Feb 3;403(6769):503-11
PMID: 10676951
-
Variations in DNA elucidate molecular networks that cause disease.
Nature. 2008 Mar 27;452(7186):429-35
PMID: 18344982
-
Embracing the complexity of genomic data for personalized medicine.
Genome Res. 2006 May;16(5):559-66
PMID: 16651662
-
Genetics of gene expression surveyed in maize, mouse and man.
Nature. 2003 Mar 20;422(6929):297-302
PMID: 12646919
-
An integrative genomics approach to infer causal associations between gene expression and disease.
Nat Genet. 2005 Jul;37(7):710-7
PMID: 15965475
-
Towards a rigorous assessment of systems biology models: the DREAM3 challenges.
PLoS One. 2010 Feb 23;5(2):e9202
PMID: 20186320
-
A modular approach for integrative analysis of large-scale gene-expression and drug-response data.
Nat Biotechnol. 2008 May;26(5):531-9
PMID: 18464786