Home LiteratureArticle Details
PMID: 21521787 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

Association studies for next-generation sequencing.

Genome research ·Vol. 21 ·No. 7 ·2011-07-00 ·Pages 1099-108

Luo L, Boerwinkle E, Xiong M

Abstract

Genome-wide association studies (GWAS) have become the primary approach for identifying genes with common variants influencing complex diseases. Despite considerable progress, the common variations identified by GWAS account for only a small fraction of disease heritability and are unlikely to explain the majority of phenotypic variations of common diseases. A potential source of the missing heritability is the contribution of rare variants. Next-generation sequencing technologies will detect millions of novel rare variants, but these technologies have three defining features: identification of a large number of rare variants, a high proportion of sequence errors, and a large proportion of missing data. These features raise challenges for testing the association of rare variants with phenotypes of interest. In this study, we use a genome continuum model and functional principal components as a general principle for developing novel and powerful association analysis methods designed for resequencing data. We use simulations to calculate the type I error rates and the power of nine alternative statistics: two functional principal component analysis (FPCA)-based statistics, the multivariate principal component analysis (MPCA)-based statistic, the weighted sum (WSS), the variable-threshold (VT) method, the generalized T(2), the collapsing method, the CMC method, and individual tests. We also examined the impact of sequence errors on their type I error rates. Finally, we apply the nine statistics to the published resequencing data set from ANGPTL4 in the Dallas Heart Study. We report that FPCA-based statistics have a higher power to detect association of rare variants and a stronger ability to filter sequence errors than the other seven methods.

MeSH Terms
Angiopoietin-Like Protein 4 Angiopoietins/genetics Computational Biology Computer Simulation Databases, Genetic Genetic Variation Genetics, Population Genome, Human Genome-Wide Association Study Genotype Humans Models, Biological Models, Statistical Multivariate Analysis Phenotype Sequence Analysis, DNA/methods,statistics & numerical data
Chemicals
ANGPTL4 protein, human Angiopoietin-Like Protein 4 Angiopoietins
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Luo Li
Human Genetics Center, University of Texas School of Public Health, Houston, TX 77030, USA.
Boerwinkle Eric
Xiong Momiao
References (27)
27 references, click to expand
  1. Rare independent mutations in renal salt handling genes contribute to blood pressure variation.
    Nat Genet. 2008 May;40(5):592-599 PMID: 18391953
  2. Potential etiologic and functional implications of genome-wide association loci for human diseases and traits.
    Proc Natl Acad Sci U S A. 2009 Jun 9;106(23):9362-7 PMID: 19474294
  3. Accurate detection and genotyping of SNPs utilizing population sequencing data.
    Genome Res. 2010 Apr;20(4):537-45 PMID: 20150320
  4. A groupwise association test for rare mutations using a weighted sum statistic.
    PLoS Genet. 2009 Feb;5(2):e1000384 PMID: 19214210
  5. To identify associations with rare variants, just WHaIT: Weighted haplotype and imputation-based tests.
    Am J Hum Genet. 2010 Nov 12;87(5):728-35 PMID: 21055717
  6. Evaluation of next generation sequencing platforms for population targeted sequencing studies.
    Genome Biol. 2009;10(3):R32 PMID: 19327155
  7. Generating samples under a Wright-Fisher neutral model of genetic variation.
    Bioinformatics. 2002 Feb;18(2):337-8 PMID: 11847089
  8. Multiple rare variants in NPC1L1 associated with reduced sterol absorption and plasma low-density lipoprotein levels.
    Proc Natl Acad Sci U S A. 2006 Feb 7;103(6):1810-5 PMID: 16449388
  9. Rare variants create synthetic genome-wide associations.
    PLoS Biol. 2010 Jan 26;8(1):e1000294 PMID: 20126254
  10. Rare variants of IFIH1, a gene implicated in antiviral responses, protect against type 1 diabetes.
    Science. 2009 Apr 17;324(5925):387-9 PMID: 19264985
  11. A comparison of bayesian methods for haplotype reconstruction from population genotype data.
    Am J Hum Genet. 2003 Nov;73(5):1162-9 PMID: 14574645
  12. Methods for detecting associations with rare variants for common diseases: application to analysis of sequence data.
    Am J Hum Genet. 2008 Sep;83(3):311-21 PMID: 18691683
  13. Common vs. rare allele hypotheses for complex diseases.
    Curr Opin Genet Dev. 2009 Jun;19(3):212-9 PMID: 19481926
  14. The prevalence of folate-remedial MTHFR enzyme variants in humans.
    Proc Natl Acad Sci U S A. 2008 Jun 10;105(23):8055-60 PMID: 18523009
  15. Detecting rare variants for complex traits using family and unrelated data.
    Genet Epidemiol. 2010 Feb;34(2):171-87 PMID: 19847924
  16. Accounting for bias from sequencing error in population genetic estimates.
    Mol Biol Evol. 2008 Jan;25(1):199-206 PMID: 17981928
  17. Pooled association tests for rare variants in exon-resequencing studies.
    Am J Hum Genet. 2010 Jun 11;86(6):832-8 PMID: 20471002
  18. Population-based resequencing of ANGPTL4 uncovers variations that reduce triglycerides and increase HDL.
    Nat Genet. 2007 Apr;39(4):513-6 PMID: 17322881
  19. Statistical analysis strategies for association studies involving rare variants.
    Nat Rev Genet. 2010 Nov;11(11):773-85 PMID: 20940738
  20. Shifting paradigm of association studies: value of rare single-nucleotide polymorphisms.
    Am J Hum Genet. 2008 Jan;82(1):100-12 PMID: 18179889
  21. The probability distribution of the amount of an individual's genome surviving to the following generation.
    Genetics. 1996 Jun;143(2):1043-9 PMID: 8725249
  22. Estimation of allele frequencies from high-coverage genome-sequencing projects.
    Genetics. 2009 May;182(1):295-301 PMID: 19293142
  23. Finding the missing heritability of complex diseases.
    Nature. 2009 Oct 8;461(7265):747-53 PMID: 19812666
  24. Population genetic inference from genomic sequence variation.
    Genome Res. 2010 Mar;20(3):291-300 PMID: 20067940
  25. The distribution of rare alleles.
    J Math Biol. 1995;33(6):602-18 PMID: 7608640
  26. De novo fragment assembly with short mate-paired reads: Does the read length matter?
    Genome Res. 2009 Feb;19(2):336-46 PMID: 19056694
  27. Human genetic variation and its contribution to complex traits.
    Nat Rev Genet. 2009 Apr;10(4):241-51 PMID: 19293820
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1549-5469
Published
2011-07-00
Epub
2011-00-26
Pages
1099-108
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC3129252
Subset
IM
Grants
NIAMS NIH HHS · P01 AR052915-01A1 · United States
NIAMS NIH HHS · P50 AR054144 · United States
NIAMS NIH HHS · 1R01AR057120-01 · United States
NHLBI NIH HHS · R01 HL106034 · United States
NHLBI NIH HHS · 1R01HL106034-01 · United States
NIAMS NIH HHS · R01 AR057120 · United States
NIAMS NIH HHS · P50 AR054144-01 · United States
NIAMS NIH HHS · P01 AR052915 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]