Home LiteratureArticle Details
PMID: 18793467 Published · epublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

Exhaustive prediction of disease susceptibility to coding base changes in the human genome.

BMC bioinformatics ·Vol. 9 Suppl 9 ·2008-08-12 ·Pages S3

Kulkarni V, Errami M, Barber R, Garner HR

Abstract

Single Nucleotide Polymorphisms (SNPs) are the most abundant form of genomic variation and can cause phenotypic differences between individuals, including diseases. Bases are subject to various levels of selection pressure, reflected in their inter-species conservation. We propose a method that is not dependant on transcription information to score each coding base in the human genome reflecting the disease probability associated with its mutation. Twelve factors likely to be associated with disease alleles were chosen as the input for a support vector machine prediction algorithm. The analysis yielded 83% sensitivity and 84% specificity in segregating disease like alleles as found in the Human Gene Mutation Database from non-disease like alleles as found in the Database of Single Nucleotide Polymorphisms. This algorithm was subsequently applied to each base within all known human genes, exhaustively confirming that interspecies conservation is the strongest factor for disease association. For each gene, the length normalized average disease potential score was calculated. Out of the 30 genes with the highest scores, 21 are directly associated with a disease. In contrast, out of the 30 genes with the lowest scores, only one is associated with a disease as found in published literature. The results strongly suggest that the highest scoring genes are enriched for those that might contribute to disease, if mutated. This method provides valuable information to researchers to identify sensitive positions in genes that have a high disease probability, enabling them to optimize experimental designs and interpret data emerging from genetic and epidemiological studies.

MeSH Terms
Algorithms Base Sequence Chromosome Mapping/methods DNA Mutational Analysis/methods Genetic Predisposition to Disease/genetics Genetic Testing/methods Genome, Human/genetics Humans Molecular Sequence Data Open Reading Frames/genetics Polymorphism, Single Nucleotide/genetics Quantitative Trait, Heritable Sequence Analysis, DNA/methods
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Kulkarni Vinayak
Mc Dermott Center for Human Growth and Development, UT Southwestern Medical Center, Dallas, TX, USA. [email protected]
Errami Mounir
Barber Robert
Garner Harold R
References (33)
33 references, click to expand
  1. Prediction of deleterious human alleles.
    Hum Mol Genet. 2001 Mar 15;10(6):591-7 PMID: 11230178
  2. TAK1 is activated in the myocardium after pressure overload and is sufficient to provoke heart failure in transgenic mice.
    Nat Med. 2000 May;6(5):556-63 PMID: 10802712
  3. A DNA polymorphism discovery resource for research on human genetic variation.
    Genome Res. 1998 Dec;8(12):1229-31 PMID: 9872978
  4. Multiple sequence alignment with the Clustal series of programs.
    Nucleic Acids Res. 2003 Jul 1;31(13):3497-500 PMID: 12824352
  5. Predicting changes in the stability of proteins and protein complexes: a study of more than 1000 mutations.
    J Mol Biol. 2002 Jul 5;320(2):369-87 PMID: 12079393
  6. Human Gene Mutation Database.
    Hum Genet. 1996 Nov;98(5):629 PMID: 8882888
  7. Selective pressures at a codon-level predict deleterious mutations in human disease genes.
    J Mol Biol. 2006 May 19;358(5):1390-404 PMID: 16584746
  8. Sequence variation in G-protein-coupled receptors: analysis of single nucleotide polymorphisms.
    Nucleic Acids Res. 2005 Mar 22;33(5):1710-21 PMID: 15784611
  9. Predicting deleterious amino acid substitutions.
    Genome Res. 2001 May;11(5):863-74 PMID: 11337480
  10. Amino acid substitution matrices from protein blocks.
    Proc Natl Acad Sci U S A. 1992 Nov 15;89(22):10915-9 PMID: 1438297
  11. PHAT: a transmembrane-specific substitution matrix. Predicted hydrophobic and transmembrane.
    Bioinformatics. 2000 Sep;16(9):760-6 PMID: 11108698
  12. A functional analysis of disease-associated mutations in the androgen receptor gene.
    Nucleic Acids Res. 2003 Apr 15;31(8):e42 PMID: 12682377
  13. The KEGG databases at GenomeNet.
    Nucleic Acids Res. 2002 Jan 1;30(1):42-6 PMID: 11752249
  14. A "silent" polymorphism in the MDR1 gene changes substrate specificity.
    Science. 2007 Jan 26;315(5811):525-8 PMID: 17185560
  15. SIFT: Predicting amino acid changes that affect protein function.
    Nucleic Acids Res. 2003 Jul 1;31(13):3812-4 PMID: 12824425
  16. Understanding human disease mutations through the use of interspecific genetic variation.
    Hum Mol Genet. 2001 Oct 1;10(21):2319-28 PMID: 11689479
  17. The comparative toxicogenomics database: a cross-species resource for building chemical-gene interaction networks.
    Toxicol Sci. 2006 Aug;92(2):587-95 PMID: 16675512
  18. NCBI's LocusLink and RefSeq.
    Nucleic Acids Res. 2000 Jan 1;28(1):126-8 PMID: 10592200
  19. As consortium plans free SNP map of human genome.
    Nature. 1999 Apr 15;398(6728):545-6 PMID: 10217129
  20. dbSNP: the NCBI database of genetic variation.
    Nucleic Acids Res. 2001 Jan 1;29(1):308-11 PMID: 11125122
  21. Evaluation of structural and evolutionary contributions to deleterious mutation prediction.
    J Mol Biol. 2002 Sep 27;322(4):891-901 PMID: 12270722
  22. The structure of haplotype blocks in the human genome.
    Science. 2002 Jun 21;296(5576):2225-9 PMID: 12029063
  23. Understanding missense mutations in the BRCA1 gene: an evolutionary approach.
    Proc Natl Acad Sci U S A. 2003 Feb 4;100(3):1151-6 PMID: 12531920
  24. Characterization of single-nucleotide polymorphisms in coding regions of human genes.
    Nat Genet. 1999 Jul;22(3):231-8 PMID: 10391209
  25. Patterns of single-nucleotide polymorphisms in candidate genes for blood-pressure homeostasis.
    Nat Genet. 1999 Jul;22(3):239-47 PMID: 10391210
  26. The essence of SNPs.
    Gene. 1999 Jul 8;234(2):177-86 PMID: 10395891
  27. SOURCE: a unified genomic resource of functional annotations, ontologies, and gene expression data.
    Nucleic Acids Res. 2003 Jan 1;31(1):219-23 PMID: 12519986
  28. Characterization of disease-associated single amino acid polymorphisms in terms of sequence and structure properties.
    J Mol Biol. 2002 Jan 25;315(4):771-86 PMID: 11812146
  29. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  30. Human Gene Mutation Database (HGMD): 2003 update.
    Hum Mutat. 2003 Jun;21(6):577-81 PMID: 12754702
  31. Identification and characterization of multi-species conserved sequences.
    Genome Res. 2003 Dec;13(12):2507-18 PMID: 14656959
  32. Predicting the functional consequences of non-synonymous single nucleotide polymorphisms: structure-based assessment of amino acid variation.
    J Mol Biol. 2001 Mar 23;307(2):683-706 PMID: 11254390
  33. Non-coding RNAs: the architects of eukaryotic complexity.
    EMBO Rep. 2001 Nov;2(11):986-91 PMID: 11713189
Article Info
Journal
BMC bioinformatics
Abbr.
BMC Bioinformatics
ISSN
1471-2105
Published
2008-08-12
Epub
2008-00-12
Pages
S3
Language
English
Region
England
NLM ID
100965194
PMCID
PMC2537574
Subset
IM
Grants
NCI NIH HHS · P50 CA070907 · United States
NCI NIH HHS · 50CA70907 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]