Home LiteratureArticle Details
PMID: 19465376 Published · ppublish English Journal Article Research Support, N.I.H., Extramural

ToppGene Suite for gene list enrichment analysis and candidate gene prioritization.

Nucleic acids research ·Vol. 37 ·No. Web Server issue ·2009-07-00 ·Pages W305-11

Chen J, Bardes EE, Aronow BJ, Jegga AG

Abstract

ToppGene Suite (http://toppgene.cchmc.org; this web site is free and open to all users and does not require a login to access) is a one-stop portal for (i) gene list functional enrichment, (ii) candidate gene prioritization using either functional annotations or network analysis and (iii) identification and prioritization of novel disease candidate genes in the interactome. Functional annotation-based disease candidate gene prioritization uses a fuzzy-based similarity measure to compute the similarity between any two genes based on semantic annotations. The similarity scores from individual features are combined into an overall score using statistical meta-analysis. A P-value of each annotation of a test gene is derived by random sampling of the whole genome. The protein-protein interaction network (PPIN)-based disease candidate gene prioritization uses social and Web networks analysis algorithms (extended versions of the PageRank and HITS algorithms, and the K-Step Markov method). We demonstrate the utility of ToppGene Suite using 20 recently reported GWAS-based gene-disease associations (including novel disease genes) representing five diseases. ToppGene ranked 19 of 20 (95%) candidate genes within the top 20%, while ToppNet ranked 12 of 16 (75%) candidate genes among the top 20%.

MeSH Terms
Animals Disease/genetics Genes Humans Internet Mice Protein Interaction Mapping Proteins/genetics Software
Chemicals
Proteins
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Chen Jing
Department of Environmental Health, University of Cincinnati, Cincinnati, OH, USA.
Bardes Eric E
Aronow Bruce J
Jegga Anil G
References (41)
41 references, click to expand
  1. Network-based global inference of human disease genes.
    Mol Syst Biol. 2008;4:189 PMID: 18463613
  2. Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
    Nucleic Acids Res. 2006 Jun 06;34(10):3067-81 PMID: 16757574
  3. A human protein-protein interaction network: a resource for annotating the proteome.
    Cell. 2005 Sep 23;122(6):957-68 PMID: 16169070
  4. BIND: the Biomolecular Interaction Network Database.
    Nucleic Acids Res. 2003 Jan 1;31(1):248-50 PMID: 12519993
  5. Genome-wide association defines more than 30 distinct susceptibility loci for Crohn's disease.
    Nat Genet. 2008 Aug;40(8):955-62 PMID: 18587394
  6. Analysis of protein sequence and interaction data for candidate disease gene prediction.
    Nucleic Acids Res. 2006;34(19):e130 PMID: 17020920
  7. Human disease genes: patterns and predictions.
    Gene. 2003 Oct 30;318:169-75 PMID: 14585509
  8. Mutation of the gene encoding the ROR2 tyrosine kinase causes autosomal recessive Robinow syndrome.
    Nat Genet. 2000 Aug;25(4):423-6 PMID: 10932187
  9. Speeding disease gene discovery by sequence based candidate prioritization.
    BMC Bioinformatics. 2005 Mar 14;6:55 PMID: 15766383
  10. Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
    Nucleic Acids Res. 2005 Mar 14;33(5):1544-52 PMID: 15767279
  11. Improved human disease candidate gene prioritization using mouse phenotype.
    BMC Bioinformatics. 2007 Oct 16;8:392 PMID: 17939863
  12. Towards a proteome-scale map of the human protein-protein interaction network.
    Nature. 2005 Oct 20;437(7062):1173-8 PMID: 16189514
  13. The human disease network.
    Proc Natl Acad Sci U S A. 2007 May 22;104(21):8685-90 PMID: 17502601
  14. POCUS: mining genomic sequence annotation to predict disease genes.
    Genome Biol. 2003;4(11):R75 PMID: 14611661
  15. Human disease genes.
    Nature. 2001 Feb 15;409(6822):853-5 PMID: 11237009
  16. Discovering disease-genes by topological features in human protein-protein interaction network.
    Bioinformatics. 2006 Nov 15;22(22):2800-5 PMID: 16954137
  17. Exploration of biological network centralities with CentiBiN.
    BMC Bioinformatics. 2006 Apr 21;7:219 PMID: 16630347
  18. Convergent functional genomics of genome-wide association data for bipolar disorder: comprehensive identification of candidate genes, pathways and mechanisms.
    Am J Med Genet B Neuropsychiatr Genet. 2009 Mar 5;150B(2):155-81 PMID: 19025758
  19. Replication of signals from recent studies of Crohn's disease identifies previously unknown disease loci for ulcerative colitis.
    Nat Genet. 2008 Jun;40(6):713-5 PMID: 18438405
  20. Mining Alzheimer disease relevant proteins from integrated protein interactome data.
    Pac Symp Biocomput. 2006;:367-78 PMID: 17094253
  21. Fuzzy measures on the Gene Ontology for gene product similarity.
    IEEE/ACM Trans Comput Biol Bioinform. 2006 Jul-Sep;3(3):263-74 PMID: 17048464
  22. Gene prioritization through genomic data fusion.
    Nat Biotechnol. 2006 May;24(5):537-44 PMID: 16680138
  23. Human protein reference database as a discovery resource for proteomics.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D497-501 PMID: 14681466
  24. Genetic determinants of ulcerative colitis include the ECM1 locus and five loci implicated in Crohn's disease.
    Nat Genet. 2008 Jun;40(6):710-2 PMID: 18438406
  25. Disruption of Abcc6 in the mouse: novel insight in the pathogenesis of pseudoxanthoma elasticum.
    Hum Mol Genet. 2005 Jul 1;14(13):1763-73 PMID: 15888484
  26. Identification of candidate disease genes by integrating Gene Ontologies and protein-interaction networks: case study of primary immunodeficiencies.
    Nucleic Acids Res. 2009 Feb;37(2):622-8 PMID: 19073697
  27. Prioritization of positional candidate genes using multiple web-based software tools.
    Twin Res Hum Genet. 2007 Dec;10(6):861-70 PMID: 18179399
  28. Candidate gene identification approach: progress and challenges.
    Int J Biol Sci. 2007 Oct 25;3(7):420-7 PMID: 17998950
  29. Common variants in the NLRP3 region contribute to Crohn's disease susceptibility.
    Nat Genet. 2009 Jan;41(1):71-6 PMID: 19098911
  30. A similarity-based method for genome-wide prediction of disease-relevant human genes.
    Bioinformatics. 2002;18 Suppl 2:S110-5 PMID: 12385992
  31. Disease candidate gene identification and prioritization using protein interaction networks.
    BMC Bioinformatics. 2009 Feb 27;10:73 PMID: 19245720
  32. Genes2Networks: connecting lists of gene symbols using mammalian protein interactions databases.
    BMC Bioinformatics. 2007 Oct 04;8:372 PMID: 17916244
  33. Murine genetic models of human disease.
    Curr Opin Genet Dev. 1994 Jun;4(3):453-60 PMID: 7919924
  34. ENDEAVOUR update: a web resource for gene prioritization in multiple species.
    Nucleic Acids Res. 2008 Jul 1;36(Web Server issue):W377-84 PMID: 18508807
  35. Replication and extension of genome-wide association study results for obesity in 4923 adults from northern Sweden.
    Hum Mol Genet. 2009 Apr 15;18(8):1489-96 PMID: 19164386
  36. Walking the interactome for prioritization of candidate disease genes.
    Am J Hum Genet. 2008 Apr;82(4):949-58 PMID: 18371930
  37. SUSPECTS: enabling fast and effective prioritization of positional candidates.
    Bioinformatics. 2006 Mar 15;22(6):773-4 PMID: 16423925
  38. A common MYBPC3 (cardiac myosin binding protein C) variant associated with cardiomyopathies in South Asia.
    Nat Genet. 2009 Feb;41(2):187-91 PMID: 19151713
  39. Protein interactions and disease: computational approaches to uncover the etiology of diseases.
    Brief Bioinform. 2007 Sep;8(5):333-46 PMID: 17638813
  40. Newly identified genetic risk variants for celiac disease related to the immune response.
    Nat Genet. 2008 Apr;40(4):395-402 PMID: 18311140
  41. The BioGRID Interaction Database: 2008 update.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D637-40 PMID: 18000002
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2009-07-00
Epub
2009-00-22
Pages
W305-11
Language
English
Region
England
NLM ID
0411011
PMCID
PMC2703978
Subset
IM
Grants
NIDDK NIH HHS · 1U01 DK70219 · United States
NIDDK NIH HHS · P30 DK078392 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]