Abstract
High-throughput DNA sequencing platforms have become widely available. As a result, personal genomes are increasingly being sequenced in research and clinical settings. However, the resulting massive amounts of variants data pose significant challenges to the average biologists and clinicians without bioinformatics skills. We developed a web server called wANNOVAR to address the critical needs for functional annotation of genetic variants from personal genomes. The server provides simple and intuitive interface to help users determine the functional significance of variants. These include annotating single nucleotide variants and insertions/deletions for their effects on genes, reporting their conservation levels (such as PhyloP and GERP++ scores), calculating their predicted functional importance scores (such as SIFT and PolyPhen scores), retrieving allele frequencies in public databases (such as the 1000 Genomes Project and NHLBI-ESP 5400 exomes), and implementing a 'variants reduction' protocol to identify a subset of potentially deleterious variants/genes. We illustrated how wANNOVAR can help draw biological insights from sequencing data, by analysing genetic variants generated on two Mendelian diseases. We conclude that wANNOVAR will help biologists and clinicians take advantage of the personal genome information to expedite scientific discoveries. The wANNOVAR server is available at http://wannovar.usc.edu, and will be continuously updated to reflect the latest annotation information.
MeSH Terms
Abnormalities, Multiple/genetics,pathology
Databases, Genetic
Exome
Genetic Variation
Genome, Human
Genomics/methods
High-Throughput Nucleotide Sequencing/methods
Humans
Internet
Limb Deformities, Congenital/genetics,pathology
Male
Mandibulofacial Dysostosis/genetics,pathology
Micrognathism/genetics,pathology
Sequence Analysis, DNA/methods
Software
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Chang Xiao
Zilkha Neurogenetic Institute, Keck School of Medicine, University of Southern California, Los Angeles, CA, USA.
Wang Kai
Supplementary Concepts
Genee-Wiedemann syndrome (Disease)
References (29)
29 references, click to expand
-
Fast and accurate short read alignment with Burrows-Wheeler transform.
Bioinformatics. 2009 Jul 15;25(14):1754-60
PMID: 19451168
-
The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data.
Genome Res. 2010 Sep;20(9):1297-303
PMID: 20644199
-
Evolutionarily conserved elements in vertebrate, insect, worm, and yeast genomes.
Genome Res. 2005 Aug;15(8):1034-50
PMID: 16024819
-
Next-generation DNA sequencing.
Nat Biotechnol. 2008 Oct;26(10):1135-45
PMID: 18846087
-
MutationTaster evaluates disease-causing potential of sequence alterations.
Nat Methods. 2010 Aug;7(8):575-6
PMID: 20676075
-
Solving the riddle of codon usage preferences: a test for translational selection.
Nucleic Acids Res. 2004 Sep 24;32(17):5036-44
PMID: 15448185
-
Accurate whole human genome sequencing using reversible terminator chemistry.
Nature. 2008 Nov 6;456(7218):53-9
PMID: 18987734
-
Segmental duplications: organization and impact within the current human genome project assembly.
Genome Res. 2001 Jun;11(6):1005-17
PMID: 11381028
-
Mapping and analysis of chromatin state dynamics in nine human cell types.
Nature. 2011 May 5;473(7345):43-9
PMID: 21441907
-
ENCODE whole-genome data in the UCSC genome browser (2011 update).
Nucleic Acids Res. 2011 Jan;39(Database issue):D871-5
PMID: 21037257
-
A map of human genome variation from population-scale sequencing.
Nature. 2010 Oct 28;467(7319):1061-73
PMID: 20981092
-
The UCSC Known Genes.
Bioinformatics. 2006 May 1;22(9):1036-46
PMID: 16500937
-
GENCODE: producing a reference annotation for ENCODE.
Genome Biol. 2006;7 Suppl 1:S4.1-9
PMID: 16925838
-
NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
Nucleic Acids Res. 2007 Jan;35(Database issue):D61-5
PMID: 17130148
-
Identifying a high fraction of the human genome to be under selective constraint using GERP++.
PLoS Comput Biol. 2010 Dec 02;6(12):e1001025
PMID: 21152010
-
TRANSFAC and its module TRANSCompel: transcriptional gene regulation in eukaryotes.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D108-10
PMID: 16381825
-
dbSNP: the NCBI database of genetic variation.
Nucleic Acids Res. 2001 Jan 1;29(1):308-11
PMID: 11125122
-
Identification of deleterious mutations within three human genomes.
Genome Res. 2009 Sep;19(9):1553-61
PMID: 19602639
-
dbNSFP: a lightweight database of human nonsynonymous SNPs and their functional predictions.
Hum Mutat. 2011 Aug;32(8):894-9
PMID: 21520341
-
Discriminative prediction of mammalian enhancers from DNA sequence.
Genome Res. 2011 Dec;21(12):2167-80
PMID: 21875935
-
ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data.
Nucleic Acids Res. 2010 Sep;38(16):e164
PMID: 20601685
-
Galaxy: a comprehensive approach for supporting accessible, reproducible, and transparent computational research in the life sciences.
Genome Biol. 2010;11(8):R86
PMID: 20738864
-
The Ensembl automatic gene annotation system.
Genome Res. 2004 May;14(5):942-50
PMID: 15123590
-
A method and server for predicting damaging missense mutations.
Nat Methods. 2010 Apr;7(4):248-9
PMID: 20354512
-
UNAFold: software for nucleic acid folding and hybridization.
Methods Mol Biol. 2008;453:3-31
PMID: 18712296
-
Exome sequencing identifies the cause of a mendelian disorder.
Nat Genet. 2010 Jan;42(1):30-5
PMID: 19915526
-
Identification of transcription factor binding sites in the human genome sequence.
Mamm Genome. 2002 Sep;13(9):510-4
PMID: 12370781
-
Using VAAST to identify an X-linked disorder resulting in lethality in male infants due to N-terminal acetyltransferase deficiency.
Am J Hum Genet. 2011 Jul 15;89(1):28-43
PMID: 21700266
-
Predicting the effects of coding non-synonymous variants on protein function using the SIFT algorithm.
Nat Protoc. 2009;4(7):1073-81
PMID: 19561590