Home LiteratureArticle Details
PMID: 15588316 Published · epublish English Journal Article

Screening large-scale association study data: exploiting interactions using random forests.

BMC genetics ·Vol. 5 ·2004-12-10 ·Pages 32

Lunetta KL, Hayward LB, Segal J, Van Eerdewegh P

Abstract

Genome-wide association studies for complex diseases will produce genotypes on hundreds of thousands of single nucleotide polymorphisms (SNPs). A logical first approach to dealing with massive numbers of SNPs is to use some test to screen the SNPs, retaining only those that meet some criterion for further study. For example, SNPs can be ranked by p-value, and those with the lowest p-values retained. When SNPs have large interaction effects but small marginal effects in a population, they are unlikely to be retained when univariate tests are used for screening. However, model-based screens that pre-specify interactions are impractical for data sets with thousands of SNPs. Random forest analysis is an alternative method that produces a single measure of importance for each predictor variable that takes into account interactions among variables without requiring model specification. Interactions increase the importance for the individual interacting variables, making them more likely to be given high importance relative to other variables. We test the performance of random forests as a screening procedure to identify small numbers of risk-associated SNPs from among large numbers of unassociated SNPs using complex disease models with up to 32 loci, incorporating both genetic heterogeneity and multi-locus interaction. Keeping other factors constant, if risk SNPs interact, the random forest importance measure significantly outperforms the Fisher Exact test as a screening tool. As the number of interacting SNPs increases, the improvement in performance of random forest analysis relative to Fisher Exact test for screening also increases. Random forests perform similarly to the univariate Fisher Exact test as a screening tool when SNPs in the analysis do not interact. In the context of large-scale genetic association studies where unknown interactions exist among true risk-associated SNPs or SNPs and environmental covariates, screening SNPs using random forest analyses can significantly reduce the number of SNPs that need to be retained for further study compared to standard univariate screening methods.

MeSH Terms
Case-Control Studies Classification Family Health Genetic Diseases, Inborn Genomics/methods Humans Linkage Disequilibrium Models, Genetic Odds Ratio Polymorphism, Single Nucleotide Siblings
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Lunetta Kathryn L
Oscient Pharmaceuticals, Inc, (formerly Genome Therapeutics Corporation), Waltham, Massachusetts, USA. [email protected]
Hayward L Brooke
Segal Jonathan
Van Eerdewegh Paul
References (21)
21 references, click to expand
  1. Using recursive partitioning for exploration and follow-up of linkage and association analyses.
    Genet Epidemiol. 1999;17 Suppl 1:S391-6 PMID: 10597468
  2. Sequence analysis using logic regression.
    Genet Epidemiol. 2001;21 Suppl 1:S626-31 PMID: 11793751
  3. Use of classification trees for association studies.
    Genet Epidemiol. 2000 Dec;19(4):323-32 PMID: 11108642
  4. Tree-based recursive partitioning methods for subdividing sibpairs into relatively more homogeneous subgroups.
    Genet Epidemiol. 2001 Apr;20(3):293-306 PMID: 11255239
  5. A combinatorial partitioning method to identify multilocus genotypic partitions that predict quantitative trait variation.
    Genome Res. 2001 Mar;11(3):458-70 PMID: 11230170
  6. Tree-based linkage and association analyses of asthma.
    Genet Epidemiol. 2001;21 Suppl 1:S317-22 PMID: 11793691
  7. Multifactor-dimensionality reduction reveals high-order interactions among estrogen-metabolism genes in sporadic breast cancer.
    Am J Hum Genet. 2001 Jul;69(1):138-47 PMID: 11404819
  8. Using data mining to address heterogeneity in the Southampton data.
    Genet Epidemiol. 2001;21 Suppl 1:S180-5 PMID: 11793666
  9. Common disease analysis using Multivariate Adaptive Regression Splines (MARS): Genetic Analysis Workshop 12 simulated sequence data.
    Genet Epidemiol. 2001;21 Suppl 1:S649-54 PMID: 11793755
  10. Founder BRCA1 and BRCA2 mutations in Ashkenazi Jews in Israel: frequency and differential penetrance in ovarian cancer and in breast-ovarian cancer families.
    Am J Hum Genet. 1997 May;60(5):1059-67 PMID: 9150153
  11. Multifactor dimensionality reduction software for detecting gene-gene and gene-environment interactions.
    Bioinformatics. 2003 Feb 12;19(3):376-82 PMID: 12584123
  12. Stochastic search variable selection for identifying multiple quantitative trait loci.
    Genetics. 2003 Jul;164(3):1129-38 PMID: 12871920
  13. A method for evaluating the results of Bayesian model selection: application to linkage analyses of attributes determined by two or more genes.
    Hum Hered. 2003;55(2-3):147-52 PMID: 12931054
  14. Mapping complex traits using Random Forests.
    BMC Genet. 2003;4 Suppl 1:S64 PMID: 14975132
  15. Use of tree-based models to identify subgroups and increase power to detect linkage to cardiovascular disease traits.
    BMC Genet. 2003;4 Suppl 1:S66 PMID: 14975134
  16. Locating disease genes using Bayesian variable selection with the Haseman-Elston method.
    BMC Genet. 2003;4 Suppl 1:S69 PMID: 14975137
  17. Tree and spline based association analysis of gene-gene interaction models for ischemic stroke.
    Stat Med. 2004 May 15;23(9):1439-53 PMID: 15116352
  18. A pilot study on the application of statistical classification procedures to molecular epidemiological data.
    Toxicol Lett. 2004 Jun 15;151(1):291-9 PMID: 15177665
  19. Linkage strategies for genetically complex traits. II. The power of affected relative pairs.
    Am J Hum Genet. 1990 Feb;46(2):229-41 PMID: 2301393
  20. Power of multifactor dimensionality reduction for detecting gene-gene interactions in the presence of genotyping error, missing data, phenocopy, and genetic heterogeneity.
    Genet Epidemiol. 2003 Feb;24(2):150-7 PMID: 12548676
  21. Classification methods for confronting heterogeneity.
    Adv Genet. 2001;42:273-86 PMID: 11037327
Article Info
Journal
BMC genetics
Abbr.
BMC Genet
ISSN
1471-2156
Published
2004-12-10
Epub
2004-00-10
Pages
32
Language
English
Region
England
NLM ID
100966978
PMCID
PMC545646
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]