Abstract
With the advance of sequencing technologies, whole exome sequencing has increasingly been used to identify mutations that cause human diseases, especially rare Mendelian diseases. Among the analysis steps, functional prediction (of being deleterious) plays an important role in filtering or prioritizing nonsynonymous SNP (NS) for further analysis. Unfortunately, different prediction algorithms use different information and each has its own strength and weakness. It has been suggested that investigators should use predictions from multiple algorithms instead of relying on a single one. However, querying predictions from different databases/Web-servers for different algorithms is both tedious and time consuming, especially when dealing with a huge number of NSs identified by exome sequencing. To facilitate the process, we developed dbNSFP (database for nonsynonymous SNPs' functional predictions). It compiles prediction scores from four new and popular algorithms (SIFT, Polyphen2, LRT, and MutationTaster), along with a conservation score (PhyloP) and other related information, for every potential NS in the human genome (a total of 75,931,005). It is the first integrated database of functional predictions from multiple algorithms for the comprehensive collection of human NSs. dbNSFP is freely available for download at http://sites.google.com/site/jpopgen/dbNSFP.
MeSH Terms
Algorithms
Computational Biology
Databases, Nucleic Acid
Genetic Association Studies
Humans
Internet
Polymorphism, Single Nucleotide/genetics
Software
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Liu Xiaoming
Human Genetics Center, School of Public Health, The University of Texas Health Science Center at Houston, Houston, Texas 77030, USA.
[email protected]
Jian Xueqiu
Boerwinkle Eric
References (28)
28 references, click to expand
-
pfSNP: An integrated potentially functional SNP resource that facilitates hypotheses generation through knowledge syntheses.
Hum Mutat. 2011 Jan;32(1):19-24
PMID: 20672376
-
Single-nucleotide evolutionary constraint scores highlight disease-causing mutations.
Nat Methods. 2010 Apr;7(4):250-1
PMID: 20354513
-
A Bayesian missing value estimation method for gene expression profile data.
Bioinformatics. 2003 Nov 1;19(16):2088-96
PMID: 14594714
-
SNPit: a federated data integration system for the purpose of functional SNP annotation.
Comput Methods Programs Biomed. 2009 Aug;95(2):181-9
PMID: 19327864
-
MutationTaster evaluates disease-causing potential of sequence alterations.
Nat Methods. 2010 Aug;7(8):575-6
PMID: 20676075
-
Pathogenic or not? And if so, then how? Studying the effects of missense mutations using bioinformatics methods.
Hum Mutat. 2009 May;30(5):703-14
PMID: 19267389
-
Predicting the effects of amino acid substitutions on protein function.
Annu Rev Genomics Hum Genet. 2006;7:61-80
PMID: 16824020
-
A method and server for predicting damaging missense mutations.
Nat Methods. 2010 Apr;7(4):248-9
PMID: 20354512
-
The Human Gene Mutation Database: 2008 update.
Genome Med. 2009 Jan 22;1(1):13
PMID: 19348700
-
Exome sequencing identifies MLL2 mutations as a cause of Kabuki syndrome.
Nat Genet. 2010 Sep;42(9):790-3
PMID: 20711175
-
Exome sequencing identifies the cause of a mendelian disorder.
Nat Genet. 2010 Jan;42(1):30-5
PMID: 19915526
-
SNPLogic: an interactive single nucleotide polymorphism selection, annotation, and prioritization system.
Nucleic Acids Res. 2009 Jan;37(Database issue):D803-9
PMID: 18984625
-
PANTHER: a library of protein families and subfamilies indexed by function.
Genome Res. 2003 Sep;13(9):2129-41
PMID: 12952881
-
Detection of nonneutral substitution rates on mammalian phylogenies.
Genome Res. 2010 Jan;20(1):110-21
PMID: 19858363
-
Predicting the effects of coding non-synonymous variants on protein function using the SIFT algorithm.
Nat Protoc. 2009;4(7):1073-81
PMID: 19561590
-
Massively parallel sequencing and rare disease.
Hum Mol Genet. 2010 Oct 15;19(R2):R119-24
PMID: 20846941
-
Pooled association tests for rare variants in exon-resequencing studies.
Am J Hum Genet. 2010 Jun 11;86(6):832-8
PMID: 20471002
-
Identification of deleterious mutations within three human genomes.
Genome Res. 2009 Sep;19(9):1553-61
PMID: 19602639
-
Dealing with missing values in large-scale studies: microarray data imputation and beyond.
Brief Bioinform. 2010 Mar;11(2):253-64
PMID: 19965979
-
Integrating common and rare genetic variation in diverse human populations.
Nature. 2010 Sep 2;467(7311):52-8
PMID: 20811451
-
The consensus coding sequence (CCDS) project: Identifying a common protein-coding gene set for the human and mouse genomes.
Genome Res. 2009 Jul;19(7):1316-23
PMID: 19498102
-
F-SNP: computationally predicted functional SNPs for disease association studies.
Nucleic Acids Res. 2008 Jan;36(Database issue):D820-4
PMID: 17986460
-
Translation efficiency in humans: tissue specificity, global optimization and differences between developmental stages.
Nucleic Acids Res. 2010 May;38(9):2964-74
PMID: 20097653
-
Bioinformatics approaches for genomics and post genomics applications of next-generation sequencing.
Brief Bioinform. 2010 Mar;11(2):181-97
PMID: 19864250
-
ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data.
Nucleic Acids Res. 2010 Sep;38(16):e164
PMID: 20601685
-
The UCSC Genome Browser database: update 2010.
Nucleic Acids Res. 2010 Jan;38(Database issue):D613-9
PMID: 19906737
-
Predicting the insurgence of human genetic diseases associated to single point protein mutations with support vector machines and evolutionary information.
Bioinformatics. 2006 Nov 15;22(22):2729-34
PMID: 16895930
-
Next generation tools for the annotation of human SNPs.
Brief Bioinform. 2009 Jan;10(1):35-52
PMID: 19181721