Home LiteratureArticle Details
PMID: 22751426 Published · epublish English Journal Article Research Support, Non-U.S. Gov't Review

Computational tools for prioritizing candidate genes: boosting disease gene discovery.

Nature reviews. Genetics ·Vol. 13 ·No. 8 ·2012-07-03 ·Pages 523-36

Moreau Y, Tranchevent LC

Abstract

At different stages of any research project, molecular biologists need to choose - often somewhat arbitrarily, even after careful statistical data analysis - which genes or proteins to investigate further experimentally and which to leave out because of limited resources. Computational methods that integrate complex, heterogeneous data sets - such as expression data, sequence information, functional annotation and the biomedical literature - allow prioritizing genes for future study in a more informed way. Such methods can substantially increase the yield of downstream studies and are becoming invaluable to researchers.

MeSH Terms
Animals Computational Biology/methods Databases, Genetic Genetic Association Studies/methods,statistics & numerical data Genetic Predisposition to Disease Haploinsufficiency/genetics Humans Mice Models, Genetic
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Moreau Yves
Department of Electrical Engineering ESAT-SCD and IBBT-KU Leuven Future Health Department, Katholieke Universiteit Leuven, Kasteelpark Arenberg 10, B-3001 Leuven, Belgium. yves.moreau@esat. kuleuven.be
Tranchevent Léon-Charles
References (122)
122 references, click to expand
  1. Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
    Nucleic Acids Res. 2006 Jun 06;34(10):3067-81 PMID: 16757574
  2. Mapping biomedical concepts onto the human genome by mining literature on chromosomal aberrations.
    Nucleic Acids Res. 2007;35(8):2533-43 PMID: 17403693
  3. Translational disease interpretation with molecular networks.
    Genome Biol. 2009;10(6):221 PMID: 19591646
  4. Improved human disease candidate gene prioritization using mouse phenotype.
    BMC Bioinformatics. 2007 Oct 16;8:392 PMID: 17939863
  5. MPV17 encodes an inner mitochondrial membrane protein and is mutated in infantile hepatic mitochondrial DNA depletion.
    Nat Genet. 2006 May;38(5):570-5 PMID: 16582910
  6. A novel gene for Usher syndrome type 2: mutations in the long isoform of whirlin are associated with retinitis pigmentosa and sensorineural hearing loss.
    Hum Genet. 2007 Apr;121(2):203-11 PMID: 17171570
  7. DECIPHER: Database of Chromosomal Imbalance and Phenotype in Humans Using Ensembl Resources.
    Am J Hum Genet. 2009 Apr;84(4):524-33 PMID: 19344873
  8. AILUN: reannotating gene expression data automatically.
    Nat Methods. 2007 Nov;4(11):879 PMID: 17971777
  9. Gene Prospector: an evidence gateway for evaluating potential susceptibility genes and interacting risk factors for human diseases.
    BMC Bioinformatics. 2008 Dec 08;9:528 PMID: 19063745
  10. The Human Phenotype Ontology: a tool for annotating and analyzing human hereditary disease.
    Am J Hum Genet. 2008 Nov;83(5):610-5 PMID: 18950739
  11. OMIM passes the 1,000-disease-gene mark.
    Nat Genet. 2000 May;25(1):11 PMID: 10802643
  12. Web tools for the prioritization of candidate disease genes.
    Methods Mol Biol. 2011;760:189-206 PMID: 21779998
  13. Prioritizing GWAS results: A review of statistical methods and recommendations for their application.
    Am J Hum Genet. 2010 Jan;86(1):6-22 PMID: 20074509
  14. MADGene: retrieval and processing of gene identifier lists for the analysis of heterogeneous microarray datasets.
    Bioinformatics. 2011 Mar 1;27(5):725-6 PMID: 21216776
  15. Mutation of Hairy-and-Enhancer-of-Split-7 in humans causes spondylocostal dysostosis.
    Hum Mol Genet. 2008 Dec 1;17(23):3761-6 PMID: 18775957
  16. The GRID: the General Repository for Interaction Datasets.
    Genome Biol. 2003;4(3):R23 PMID: 12620108
  17. CGI: a new approach for prioritizing genes by combining gene expression and protein-protein interaction data.
    Bioinformatics. 2007 Jan 15;23(2):215-21 PMID: 17098772
  18. Exploring the human genome with functional maps.
    Genome Res. 2009 Jun;19(6):1093-106 PMID: 19246570
  19. CANDID: a flexible method for prioritizing candidate genes for complex human traits.
    Genet Epidemiol. 2008 Dec;32(8):779-90 PMID: 18613097
  20. Large-scale benchmark of Endeavour using MetaCore maps.
    Bioinformatics. 2010 Aug 1;26(15):1922-3 PMID: 20538729
  21. Annotating the human genome with Disease Ontology.
    BMC Genomics. 2009 Jul 07;10 Suppl 1:S6 PMID: 19594883
  22. The impact of multifunctional genes on "guilt by association" analysis.
    PLoS One. 2011 Feb 18;6(2):e17258 PMID: 21364756
  23. Severely incapacitating mutations in patients with extreme short stature identify RNA-processing endoribonuclease RMRP as an essential cell growth regulator.
    Am J Hum Genet. 2005 Nov;77(5):795-806 PMID: 16252239
  24. STRING: a web-server to retrieve and display the repeatedly occurring neighbourhood of a gene.
    Nucleic Acids Res. 2000 Sep 15;28(18):3442-4 PMID: 10982861
  25. Gene Expression Omnibus: NCBI gene expression and hybridization array data repository.
    Nucleic Acids Res. 2002 Jan 1;30(1):207-10 PMID: 11752295
  26. Reconstruction of a functional human gene network, with an application for prioritizing positional candidate genes.
    Am J Hum Genet. 2006 Jun;78(6):1011-25 PMID: 16685651
  27. Systematic and integrative analysis of large gene lists using DAVID bioinformatics resources.
    Nat Protoc. 2009;4(1):44-57 PMID: 19131956
  28. Association of genes to genetically inherited diseases using data mining.
    Nat Genet. 2002 Jul;31(3):316-9 PMID: 12006977
  29. A strategy to search for common obesity and type 2 diabetes genes.
    Trends Endocrinol Metab. 2007 Jan-Feb;18(1):19-26 PMID: 17126559
  30. Network medicine: a network-based approach to human disease.
    Nat Rev Genet. 2011 Jan;12(1):56-68 PMID: 21164525
  31. Annotation transfer between genomes: protein-protein interologs and protein-DNA regulogs.
    Genome Res. 2004 Jun;14(6):1107-18 PMID: 15173116
  32. Collaboratively charting the gene-to-phenotype network of human congenital heart defects.
    Genome Med. 2010 Mar 01;2(3):16 PMID: 20193066
  33. Update of the G2D tool for prioritization of gene candidates to inherited diseases.
    Nucleic Acids Res. 2007 Jul;35(Web Server issue):W212-6 PMID: 17478516
  34. Genome-wide identification of genes likely to be involved in human genetic disease.
    Nucleic Acids Res. 2004 Jun 04;32(10):3108-14 PMID: 15181176
  35. Whole-genome sequencing in a patient with Charcot-Marie-Tooth neuropathy.
    N Engl J Med. 2010 Apr 1;362(13):1181-91 PMID: 20220177
  36. Edgetic perturbation models of human inherited disorders.
    Mol Syst Biol. 2009;5:321 PMID: 19888216
  37. KEGG for integration and interpretation of large-scale molecular data sets.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D109-14 PMID: 22080510
  38. Computational approaches to disease-gene prediction: rationale, classification and successes.
    FEBS J. 2012 Mar;279(5):678-96 PMID: 22221742
  39. Genetic dissection of the biotic stress response using a genome-scale gene network for rice.
    Proc Natl Acad Sci U S A. 2011 Nov 8;108(45):18548-53 PMID: 22042862
  40. Comparison of genomic and proteomic data in recurrent airway obstruction affected horses using Ingenuity Pathway Analysis®.
    BMC Vet Res. 2011 Aug 15;7:48 PMID: 21843342
  41. A statistical framework for genomic data fusion.
    Bioinformatics. 2004 Nov 1;20(16):2626-35 PMID: 15130933
  42. GoPubMed: exploring PubMed with the Gene Ontology.
    Nucleic Acids Res. 2005 Jul 1;33(Web Server issue):W783-6 PMID: 15980585
  43. Outcome of array CGH analysis for 255 subjects with intellectual disability and search for candidate genes using bioinformatics.
    Hum Genet. 2010 Aug;128(2):179-94 PMID: 20512354
  44. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  45. "Guilt by association" is the exception rather than the rule in gene networks.
    PLoS Comput Biol. 2012;8(3):e1002444 PMID: 22479173
  46. Conceptual thinking for in silico prioritization of candidate disease genes.
    Methods Mol Biol. 2011;760:175-87 PMID: 21779997
  47. Analysis of protein sequence and interaction data for candidate disease gene prediction.
    Nucleic Acids Res. 2006;34(19):e130 PMID: 17020920
  48. Towards a cyberinfrastructure for the biological sciences: progress, visions and challenges.
    Nat Rev Genet. 2008 Sep;9(9):678-88 PMID: 18714290
  49. The genetic association database.
    Nat Genet. 2004 May;36(5):431-2 PMID: 15118671
  50. Biofilter: a knowledge-integration system for the multi-locus analysis of genome-wide association studies.
    Pac Symp Biocomput. 2009;:368-79 PMID: 19209715
  51. Identification of Parkinson's disease candidate genes using CAESAR and screening of MAPT and SNCAIP in South African Parkinson's disease patients.
    J Neural Transm (Vienna). 2011 Jun;118(6):889-97 PMID: 21344240
  52. Cytoscape: a software environment for integrated models of biomolecular interaction networks.
    Genome Res. 2003 Nov;13(11):2498-504 PMID: 14597658
  53. Prioritization of positional candidate genes using multiple web-based software tools.
    Twin Res Hum Genet. 2007 Dec;10(6):861-70 PMID: 18179399
  54. Call to work together on microarray data analysis.
    Nature. 2001 Jun 21;411(6840):885 PMID: 11418825
  55. Transcriptional profiling of bovine milk using RNA sequencing.
    BMC Genomics. 2012 Jan 25;13:45 PMID: 22276848
  56. Interactome networks and human disease.
    Cell. 2011 Mar 18;144(6):986-98 PMID: 21414488
  57. Kernel-based data fusion for gene prioritization.
    Bioinformatics. 2007 Jul 1;23(13):i125-32 PMID: 17646288
  58. Critical assessment of methods of protein structure prediction (CASP)--round IX.
    Proteins. 2011;79 Suppl 10:1-5 PMID: 21997831
  59. Human disease genes: patterns and predictions.
    Gene. 2003 Oct 30;318:169-75 PMID: 14585509
  60. TEAM: a tool for the integration of expression, and linkage and association maps.
    Eur J Hum Genet. 2004 Aug;12(8):633-8 PMID: 15114375
  61. Meta-analysis of heterogeneous data sources for genome-scale identification of risk genes in complex phenotypes.
    Genet Epidemiol. 2011 Jul;35(5):318-32 PMID: 21484861
  62. A survey of current software for network analysis in molecular biology.
    Hum Genomics. 2010 Jun;4(5):353-60 PMID: 20650822
  63. BioCreative III interactive task: an overview.
    BMC Bioinformatics. 2011 Oct 03;12 Suppl 8:S4 PMID: 22151968
  64. A new web-based data mining tool for the identification of candidate genes for human genetic disorders.
    Eur J Hum Genet. 2003 Jan;11(1):57-63 PMID: 12529706
  65. Prioritizing candidate disease genes by network-based boosting of genome-wide association data.
    Genome Res. 2011 Jul;21(7):1109-21 PMID: 21536720
  66. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles.
    Proc Natl Acad Sci U S A. 2005 Oct 25;102(43):15545-50 PMID: 16199517
  67. BioMart--biological queries made easy.
    BMC Genomics. 2009 Jan 14;10:22 PMID: 19144180
  68. The impact of incomplete knowledge on evaluation: an experimental benchmark for protein function prediction.
    Bioinformatics. 2009 Sep 15;25(18):2404-10 PMID: 19561015
  69. HHEX gene polymorphisms are associated with type 2 diabetes in the Dutch Breda cohort.
    Eur J Hum Genet. 2008 May;16(5):652-6 PMID: 18231124
  70. Génie: literature-based gene prioritization at multi genomic scale.
    Nucleic Acids Res. 2011 Jul;39(Web Server issue):W455-61 PMID: 21609954
  71. Inparanoid: a comprehensive database of eukaryotic orthologs.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D476-80 PMID: 15608241
  72. Online Mendelian Inheritance in Man (OMIM).
    Hum Mutat. 2000;15(1):57-61 PMID: 10612823
  73. A human phenome-interactome network of protein complexes implicated in genetic disorders.
    Nat Biotechnol. 2007 Mar;25(3):309-16 PMID: 17344885
  74. The STRING database in 2011: functional interaction networks of proteins, globally integrated and scored.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D561-8 PMID: 21045058
  75. Exome sequencing and disease-network analysis of a single family implicate a mutation in KIF1A in hereditary spastic paraparesis.
    Genome Res. 2011 May;21(5):658-64 PMID: 21487076
  76. Bioinformatics enrichment tools: paths toward the comprehensive functional analysis of large gene lists.
    Nucleic Acids Res. 2009 Jan;37(1):1-13 PMID: 19033363
  77. Critical assessment of methods of protein structure prediction (CASP): round II.
    Proteins. 1997;Suppl 1:2-6 PMID: 9485489
  78. GeneDistiller--distilling candidate genes from linkage intervals.
    PLoS One. 2008;3(12):e3874 PMID: 19057649
  79. PINTA: a web server for network-based gene prioritization from expression data.
    Nucleic Acids Res. 2011 Jul;39(Web Server issue):W334-8 PMID: 21602267
  80. Advances in translational bioinformatics: computational approaches for the hunting of disease genes.
    Brief Bioinform. 2010 Jan;11(1):96-110 PMID: 20007728
  81. Speeding disease gene discovery by sequence based candidate prioritization.
    BMC Bioinformatics. 2005 Mar 14;6:55 PMID: 15766383
  82. Computationally driven, quantitative experiments discover genes required for mitochondrial biogenesis.
    PLoS Genet. 2009 Mar;5(3):e1000407 PMID: 19300474
  83. Recurring mutations found by sequencing an acute myeloid leukemia genome.
    N Engl J Med. 2009 Sep 10;361(11):1058-66 PMID: 19657110
  84. Pathway mapping tools for analysis of high content data.
    Methods Mol Biol. 2007;356:319-50 PMID: 16988414
  85. Overview of BioCreAtIvE: critical assessment of information extraction for biology.
    BMC Bioinformatics. 2005;6 Suppl 1:S1 PMID: 15960821
  86. Comparison of automated candidate gene prediction systems using genes implicated in type 2 diabetes by genome-wide association studies.
    BMC Bioinformatics. 2009 Jan 30;10 Suppl 1:S69 PMID: 19208173
  87. ENDEAVOUR update: a web resource for gene prioritization in multiple species.
    Nucleic Acids Res. 2008 Jul 1;36(Web Server issue):W377-84 PMID: 18508807
  88. The biological coherence of human phenome databases.
    Am J Hum Genet. 2009 Dec;85(6):801-8 PMID: 20004759
  89. The Human Gene Mutation Database: 2008 update.
    Genome Med. 2009 Jan 22;1(1):13 PMID: 19348700
  90. In silico gene prioritization by integrating multiple data sources.
    PLoS One. 2011;6(6):e21137 PMID: 21731658
  91. Human disease: something old, something new.
    Nat Rev Genet. 2011 Jun;12(6):382-3 PMID: 21556015
  92. Ensembl 2012.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D84-90 PMID: 22086963
  93. The power of protein interaction networks for associating genes with diseases.
    Bioinformatics. 2010 Apr 15;26(8):1057-63 PMID: 20185403
  94. Walking the interactome for prioritization of candidate disease genes.
    Am J Hum Genet. 2008 Apr;82(4):949-58 PMID: 18371930
  95. Human Gene Mutation Database (HGMD): 2003 update.
    Hum Mutat. 2003 Jun;21(6):577-81 PMID: 12754702
  96. A Bayesian framework for combining heterogeneous data sources for gene function prediction (in Saccharomyces cerevisiae).
    Proc Natl Acad Sci U S A. 2003 Jul 8;100(14):8348-53 PMID: 12826619
  97. Genome-wide prioritization of disease genes and identification of disease-disease associations from an integrated human functional linkage network.
    Genome Biol. 2009;10(9):R91 PMID: 19728866
  98. Infantile cerebral and cerebellar atrophy is associated with a mutation in the MED17 subunit of the transcription preinitiation mediator complex.
    Am J Hum Genet. 2010 Nov 12;87(5):667-70 PMID: 20950787
  99. DNA microarrays: vital statistics.
    Nature. 2003 Aug 7;424(6949):610-2 PMID: 12904757
  100. The UCSC Genome Browser database: extensions and updates 2011.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D918-23 PMID: 22086951
  101. A gene-coexpression network for global discovery of conserved genetic modules.
    Science. 2003 Oct 10;302(5643):249-55 PMID: 12934013
  102. Integrating computational biology and forward genetics in Drosophila.
    PLoS Genet. 2009 Jan;5(1):e1000351 PMID: 19165344
  103. Towards a proteome-scale map of the human protein-protein interaction network.
    Nature. 2005 Oct 20;437(7062):1173-8 PMID: 16189514
  104. Fatal cardiac arrhythmia and long-QT syndrome in a new form of congenital generalized lipodystrophy with muscle rippling (CGL4) due to PTRF-CAVIN mutations.
    PLoS Genet. 2010 Mar 12;6(3):e1000874 PMID: 20300641
  105. Linking genes to literature: text mining, information extraction, and retrieval applications for biology.
    Genome Biol. 2008;9 Suppl 2:S8 PMID: 18834499
  106. GPSy: a cross-species gene prioritization system for conserved biological processes--application in male gamete development.
    Nucleic Acids Res. 2012 Jul;40(Web Server issue):W458-65 PMID: 22570409
  107. Protein interactions and disease: computational approaches to uncover the etiology of diseases.
    Brief Bioinform. 2007 Sep;8(5):333-46 PMID: 17638813
  108. G2D: a tool for mining genes associated with disease.
    BMC Genet. 2005 Aug 22;6:45 PMID: 16115313
  109. Molecular networks as sensors and drivers of common human diseases.
    Nature. 2009 Sep 10;461(7261):218-23 PMID: 19741703
  110. STITCH: interaction networks of chemicals and proteins.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D684-8 PMID: 18084021
  111. Crowdsourcing network inference: the DREAM predictive signaling network challenge.
    Sci Signal. 2011 Aug 30;4(189):mr7 PMID: 21900204
  112. A literature network of human genes for high-throughput analysis of gene expression.
    Nat Genet. 2001 May;28(1):21-8 PMID: 11326270
  113. PosMed (Positional Medline): prioritizing genes with an artificial neural network comprising medical documents to accelerate positional cloning.
    Nucleic Acids Res. 2009 Jul;37(Web Server issue):W147-52 PMID: 19468046
  114. Facts from text: can text mining help to scale-up high-quality manual curation of gene products with ontologies?
    Brief Bioinform. 2008 Nov;9(6):466-78 PMID: 19060303
  115. The human disease network.
    Proc Natl Acad Sci U S A. 2007 May 22;104(21):8685-90 PMID: 17502601
  116. PolySearch: a web-based text mining system for extracting relationships between human diseases, genes, mutations, drugs and metabolites.
    Nucleic Acids Res. 2008 Jul 1;36(Web Server issue):W399-405 PMID: 18487273
  117. Genes to diseases (G2D) computational method to identify asthma candidate genes.
    PLoS One. 2008 Aug 06;3(8):e2907 PMID: 18682798
  118. Systematic prediction of gene function in Arabidopsis thaliana using a probabilistic functional gene network.
    Nat Protoc. 2011 Aug 25;6(9):1429-42 PMID: 21886106
  119. ToppGene Suite for gene list enrichment analysis and candidate gene prioritization.
    Nucleic Acids Res. 2009 Jul;37(Web Server issue):W305-11 PMID: 19465376
  120. Needles in stacks of needles: finding disease-causal variants in a wealth of genomic data.
    Nat Rev Genet. 2011 Aug 18;12(9):628-40 PMID: 21850043
  121. ArrayExpress update--an archive of microarray and high-throughput sequencing-based functional genomics experiments.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D1002-4 PMID: 21071405
  122. Haploinsufficiency of TAB2 causes congenital heart defects in humans.
    Am J Hum Genet. 2010 Jun 11;86(6):839-49 PMID: 20493459
Article Info
Journal
Nature reviews. Genetics
Abbr.
Nat Rev Genet
ISSN
1471-0064
Published
2012-07-03
Epub
2012-00-03
Pages
523-36
Language
English
Region
England
NLM ID
100962779
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]