Home LiteratureArticle Details
PMID: 17020920 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't Validation Study

Analysis of protein sequence and interaction data for candidate disease gene prediction.

Nucleic acids research ·Vol. 34 ·No. 19 ·2006-00-00 ·Pages e130

George RA, Liu JY, Feng LL, Bryson-Richardson RJ, Fatkin D, Wouters MA

Abstract

Linkage analysis is a successful procedure to associate diseases with specific genomic regions. These regions are often large, containing hundreds of genes, which make experimental methods employed to identify the disease gene arduous and expensive. We present two methods to prioritize candidates for further experimental study: Common Pathway Scanning (CPS) and Common Module Profiling (CMP). CPS is based on the assumption that common phenotypes are associated with dysfunction in proteins that participate in the same complex or pathway. CPS applies network data derived from protein-protein interaction (PPI) and pathway databases to identify relationships between genes. CMP identifies likely candidates using a domain-dependent sequence similarity approach, based on the hypothesis that disruption of genes of similar function will lead to the same phenotype. Both algorithms use two forms of input data: known disease genes or multiple disease loci. When using known disease genes as input, our combined methods have a sensitivity of 0.52 and a specificity of 0.97 and reduce the candidate list by 13-fold. Using multiple loci, our methods successfully identify disease genes for all benchmark diseases with a sensitivity of 0.84 and a specificity of 0.63. Our combined approach prioritizes good candidates and will accelerate the disease gene discovery process.

MeSH Terms
Algorithms Computational Biology Databases, Protein Genes Genetic Predisposition to Disease Humans Phenotype Protein Interaction Mapping Protein Structure, Tertiary Proteins/genetics Sequence Analysis, Protein/methods
Chemicals
Proteins
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
George Richard A
Computational Biology & Bioinformatics Program, Sydney, NSW, Australia.
Liu Jason Y
Feng Lina L
Bryson-Richardson Robert J
Fatkin Diane
Wouters Merridee A
References (42)
42 references, click to expand
  1. Human disease genes.
    Nature. 2001 Feb 15;409(6822):853-5 PMID: 11237009
  2. Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
    Nat Genet. 2000 May;25(1):25-9 PMID: 10802651
  3. MINT: a Molecular INTeraction database.
    FEBS Lett. 2002 Feb 20;513(1):135-40 PMID: 11911893
  4. Comparative assessment of large-scale data sets of protein-protein interactions.
    Nature. 2002 May 23;417(6887):399-403 PMID: 12000970
  5. Association of genes to genetically inherited diseases using data mining.
    Nat Genet. 2002 Jul;31(3):316-9 PMID: 12006977
  6. Protein domain identification and improved sequence similarity searching using PSI-BLAST.
    Proteins. 2002 Sep 1;48(4):672-81 PMID: 12211035
  7. Beyond Mendel: an evolving view of human genetic disease transmission.
    Nat Rev Genet. 2002 Oct;3(10):779-89 PMID: 12360236
  8. A similarity-based method for genome-wide prediction of disease-relevant human genes.
    Bioinformatics. 2002;18 Suppl 2:S110-5 PMID: 12385992
  9. BIND: the Biomolecular Interaction Network Database.
    Nucleic Acids Res. 2003 Jan 1;31(1):248-50 PMID: 12519993
  10. Integration of text- and data-mining using ontologies successfully selects disease gene candidates.
    Nucleic Acids Res. 2005;33(5):1544-52 PMID: 15767279
  11. Online predicted human interaction database.
    Bioinformatics. 2005 May 1;21(9):2076-82 PMID: 15657099
  12. Consolidating the set of known human protein-protein interactions in preparation for large-scale mapping of the human interactome.
    Genome Biol. 2005;6(5):R40 PMID: 15892868
  13. WW domains provide a platform for the assembly of multiprotein networks.
    Mol Cell Biol. 2005 Aug;25(16):7092-106 PMID: 16055720
  14. Effective function annotation through catalytic residue conservation.
    Proc Natl Acad Sci U S A. 2005 Aug 30;102(35):12299-304 PMID: 16037208
  15. G2D: a tool for mining genes associated with disease.
    BMC Genet. 2005;6:45 PMID: 16115313
  16. A human protein-protein interaction network: a resource for annotating the proteome.
    Cell. 2005 Sep 23;122(6):957-68 PMID: 16169070
  17. Towards a proteome-scale map of the human protein-protein interaction network.
    Nature. 2005 Oct 20;437(7062):1173-8 PMID: 16189514
  18. Database resources of the National Center for Biotechnology Information.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D173-80 PMID: 16381840
  19. Pathguide: a pathway resource list.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D504-6 PMID: 16381921
  20. A quantitative protein interaction network for the ErbB receptors using protein microarrays.
    Nature. 2006 Jan 12;439(7073):168-74 PMID: 16273093
  21. Analysis of the human protein interactome and comparison with yeast, worm and fly interaction datasets.
    Nat Genet. 2006 Mar;38(3):285-93 PMID: 16501559
  22. Systematic identification of functional orthologs based on protein network comparison.
    Genome Res. 2006 Mar;16(3):428-35 PMID: 16510899
  23. SUSPECTS: enabling fast and effective prioritization of positional candidates.
    Bioinformatics. 2006 Mar 15;22(6):773-4 PMID: 16423925
  24. Reconstruction of a functional human gene network, with an application for prioritizing positional candidate genes.
    Am J Hum Genet. 2006 Jun;78(6):1011-25 PMID: 16685651
  25. A genome-wide association study of nonsynonymous SNPs identifies a type 1 diabetes locus in the interferon-induced helicase (IFIH1) region.
    Nat Genet. 2006 Jun;38(6):617-9 PMID: 16699517
  26. Variants in the GH-IGF axis confer susceptibility to lung cancer.
    Genome Res. 2006 Jun;16(6):693-701 PMID: 16741161
  27. Computational disease gene identification: a concert of methods prioritizes type 2 diabetes and obesity candidate genes.
    Nucleic Acids Res. 2006;34(10):3067-81 PMID: 16757574
  28. Predicting disease genes using protein-protein interactions.
    J Med Genet. 2006 Aug;43(8):691-8 PMID: 16611749
  29. The ciliopathies: an emerging class of human genetic disorders.
    Annu Rev Genomics Hum Genet. 2006;7:125-48 PMID: 16722803
  30. A new web-based data mining tool for the identification of candidate genes for human genetic disorders.
    Eur J Hum Genet. 2003 Jan;11(1):57-63 PMID: 12529706
  31. eVOC: a controlled vocabulary for unifying gene expression data.
    Genome Res. 2003 Jun;13(6A):1222-30 PMID: 12799354
  32. New methods for finding disease-susceptibility genes: impact and potential.
    Genome Biol. 2003;4(10):119 PMID: 14519189
  33. Development of human protein reference database as an initial platform for approaching systems biology in humans.
    Genome Res. 2003 Oct;13(10):2363-71 PMID: 14525934
  34. POCUS: mining genomic sequence annotation to predict disease genes.
    Genome Biol. 2003;4(11):R75 PMID: 14611661
  35. The Pfam protein families database.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D138-41 PMID: 14681378
  36. The KEGG resource for deciphering the genome.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D277-80 PMID: 14681412
  37. Identification of common molecular subsequences.
    J Mol Biol. 1981 Mar 25;147(1):195-7 PMID: 7265238
  38. Improved tools for biological sequence comparison.
    Proc Natl Acad Sci U S A. 1988 Apr;85(8):2444-8 PMID: 3162770
  39. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
  40. InterPro, progress and status in 2005.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D201-5 PMID: 15608177
  41. GenBank.
    Nucleic Acids Res. 2005 Jan 1;33(Database issue):D34-8 PMID: 15608212
  42. Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic disorders.
    Nucleic Acids Res. 2002 Jan 1;30(1):52-5 PMID: 11752252
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2006-00-00
Epub
2006-00-04
Pages
e130
Language
English
Region
England
NLM ID
0411011
PMCID
PMC1636487
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]