Home LiteratureArticle Details
PMID: 27084946 Published · ppublish English Journal Article

DanQ: a hybrid convolutional and recurrent deep neural network for quantifying the function of DNA sequences.

Nucleic acids research ·Vol. 44 ·No. 11 ·2016-00-20 ·Pages e107

Quang D, Xie X

Abstract

Modeling the properties and functions of DNA sequences is an important, but challenging task in the broad field of genomics. This task is particularly difficult for non-coding DNA, the vast majority of which is still poorly understood in terms of function. A powerful predictive model for the function of non-coding DNA can have enormous benefit for both basic science and translational research because over 98% of the human genome is non-coding and 93% of disease-associated variants lie in these regions. To address this need, we propose DanQ, a novel hybrid convolutional and bi-directional long short-term memory recurrent neural network framework for predicting non-coding function de novo from sequence. In the DanQ model, the convolution layer captures regulatory motifs, while the recurrent layer captures long-term dependencies between the motifs in order to learn a regulatory 'grammar' to improve predictions. DanQ improves considerably upon other models across several metrics. For some regulatory markers, DanQ can achieve over a 50% relative improvement in the area under the precision-recall curve metric compared to related models. We have made the source code available at the github repository http://github.com/uci-cbcl/DanQ.

MeSH Terms
DNA Genome-Wide Association Study Genomics/methods Humans Neural Networks, Computer Polymorphism, Single Nucleotide Quantitative Trait Loci ROC Curve Sequence Analysis, DNA Software Web Browser
Chemicals
DNA
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Quang Daniel
Department of Computer Science University of California, Irvine, CA 92697, USA Center for Complex Biological Systems University of California, Irvine, CA 92697, USA.
Xie Xiaohui
Department of Computer Science University of California, Irvine, CA 92697, USA Center for Complex Biological Systems University of California, Irvine, CA 92697, USA [email protected].
References (18)
18 references, click to expand
  1. JASPAR 2016: a major expansion and update of the open-access database of transcription factor binding profiles.
    Nucleic Acids Res. 2016 Jan 4;44(D1):D110-5 PMID: 26531826
  2. Integrative analysis of 111 reference human epigenomes.
    Nature. 2015 Feb 19;518(7539):317-30 PMID: 25693563
  3. RSAT 2015: Regulatory Sequence Analysis Tools.
    Nucleic Acids Res. 2015 Jul 1;43(W1):W50-6 PMID: 25904632
  4. Potential etiologic and functional implications of genome-wide association loci for human diseases and traits.
    Proc Natl Acad Sci U S A. 2009 Jun 9;106(23):9362-7 PMID: 19474294
  5. An integrated map of genetic variation from 1,092 human genomes.
    Nature. 2012 Nov 1;491(7422):56-65 PMID: 23128226
  6. The NHGRI GWAS Catalog, a curated resource of SNP-trait associations.
    Nucleic Acids Res. 2014 Jan;42(Database issue):D1001-6 PMID: 24316577
  7. DANN: a deep learning approach for annotating the pathogenicity of genetic variants.
    Bioinformatics. 2015 Mar 1;31(5):761-3 PMID: 25338716
  8. Motif signatures in stretch enhancers are enriched for disease-associated genetic variants.
    Epigenetics Chromatin. 2015 Jul 16;8:23 PMID: 26180553
  9. An integrated encyclopedia of DNA elements in the human genome.
    Nature. 2012 Sep 6;489(7414):57-74 PMID: 22955616
  10. GRASP: analysis of genotype-phenotype results from 1390 genome-wide association studies and corresponding open access database.
    Bioinformatics. 2014 Jun 15;30(12):i185-94 PMID: 24931982
  11. A method to predict the impact of regulatory variants from DNA sequence.
    Nat Genet. 2015 Aug;47(8):955-61 PMID: 26075791
  12. EXTREME: an online EM algorithm for motif discovery.
    Bioinformatics. 2014 Jun 15;30(12):1667-73 PMID: 24532725
  13. Predicting effects of noncoding variants with deep learning-based sequence model.
    Nat Methods. 2015 Oct;12(10):931-4 PMID: 26301843
  14. Predicting the sequence specificities of DNA- and RNA-binding proteins by deep learning.
    Nat Biotechnol. 2015 Aug;33(8):831-8 PMID: 26213851
  15. Framewise phoneme classification with bidirectional LSTM and other neural network architectures.
    Neural Netw. 2005 Jun-Jul;18(5-6):602-10 PMID: 16112549
  16. Deep learning.
    Nature. 2015 May 28;521(7553):436-44 PMID: 26017442
  17. Quantifying similarity between motifs.
    Genome Biol. 2007;8(2):R24 PMID: 17324271
  18. Enhanced regulatory sequence prediction using gapped k-mer features.
    PLoS Comput Biol. 2014 Jul 17;10(7):e1003711 PMID: 25033408
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2016-00-20
Epub
2016-00-15
Pages
e107
Language
English
Region
England
NLM ID
0411011
PMCID
PMC4914104
Subset
IM
Grants
NIBIB NIH HHS · T32 EB009418 · United States
NHGRI NIH HHS · R01 HG006870 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]