Home LiteratureArticle Details
PMID: 24172663 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

EXCAVATOR: detecting copy number variants from whole-exome sequencing data.

Genome biology ·Vol. 14 ·No. 10 ·2013-00-00 ·Pages R120

Magi A, Tattini L, Cifola I, D'Aurizio R, Benelli M, Mangano E, Battaglia C, Bonora E, Kurg A, Seri M, Magini P, Giusti B, Romeo G, Pippucci T, De Bellis G, Abbate R, Gensini GF

Abstract

We developed a novel software tool, EXCAVATOR, for the detection of copy number variants (CNVs) from whole-exome sequencing data. EXCAVATOR combines a three-step normalization procedure with a novel heterogeneous hidden Markov model algorithm and a calling method that classifies genomic regions into five copy number states. We validate EXCAVATOR on three datasets and compare the results with three other methods. These analyses show that EXCAVATOR outperforms the other methods and is therefore a valuable tool for the investigation of CNVs in largescale projects, as well as in clinical research and diagnostics. EXCAVATOR is freely available at http://sourceforge.net/projects/excavatortool/.

MeSH Terms
Algorithms Computational Biology/methods DNA Copy Number Variations Exome Genome Genomics High-Throughput Nucleotide Sequencing Humans Intellectual Disability/genetics Markov Chains Melanoma/genetics,pathology Polymorphism, Single Nucleotide ROC Curve Reproducibility of Results Software
Authors & Affiliations
17 authors, click to expand affiliations / ORCID
Magi Alberto
Tattini Lorenzo
Cifola Ingrid
D'Aurizio Romina
Benelli Matteo
Mangano Eleonora
Battaglia Cristina
Bonora Elena
Kurg Ants
Seri Marco
Magini Pamela
Giusti Betti
Romeo Giovanni
Pippucci Tommaso
De Bellis Gianluca
Abbate Rosanna
Gensini Gian Franco
References (45)
45 references, click to expand
  1. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data.
    Genome Res. 2010 Sep;20(9):1297-303 PMID: 20644199
  2. SOAP2: an improved ultrafast tool for short read alignment.
    Bioinformatics. 2009 Aug 1;25(15):1966-7 PMID: 19497933
  3. Sensitive and accurate detection of copy number variants using read depth of coverage.
    Genome Res. 2009 Sep;19(9):1586-92 PMID: 19657104
  4. A shifting level model algorithm that identifies aberrations in array-CGH data.
    Biostatistics. 2010 Apr;11(2):265-80 PMID: 19948744
  5. Large-scale copy number polymorphism in the human genome.
    Science. 2004 Jul 23;305(5683):525-8 PMID: 15273396
  6. Detection of structural variants and indels within exome data.
    Nat Methods. 2011 Dec 18;9(2):176-8 PMID: 22179552
  7. Circular binary segmentation for the analysis of array-based DNA copy number data.
    Biostatistics. 2004 Oct;5(4):557-72 PMID: 15475419
  8. Detecting common copy number variants in high-throughput sequencing data by using JointSLM algorithm.
    Nucleic Acids Res. 2011 May;39(10):e65 PMID: 21321017
  9. Genome structural variation discovery and genotyping.
    Nat Rev Genet. 2011 May;12(5):363-76 PMID: 21358748
  10. The uniqueome: a mappability resource for short-tag sequencing.
    Bioinformatics. 2011 Jan 15;27(2):272-4 PMID: 21075741
  11. Global variation in copy number in the human genome.
    Nature. 2006 Nov 23;444(7118):444-54 PMID: 17122850
  12. Whole-genome sequencing and variant discovery in C. elegans.
    Nat Methods. 2008 Feb;5(2):183-8 PMID: 18204455
  13. Discovery and statistical genotyping of copy-number variation from whole-exome sequencing depth.
    Am J Hum Genet. 2012 Oct 5;91(4):597-607 PMID: 23040492
  14. Read count approach for DNA copy number variants detection.
    Bioinformatics. 2012 Feb 15;28(4):470-8 PMID: 22199393
  15. Sequence and structural variation in a human genome uncovered by short-read, massively parallel ligation sequencing using two-base encoding.
    Genome Res. 2009 Sep;19(9):1527-41 PMID: 19546169
  16. Integrated detection and population-genetic analysis of SNPs and copy number variation.
    Nat Genet. 2008 Oct;40(10):1166-74 PMID: 18776908
  17. The complete genome of an individual by massively parallel DNA sequencing.
    Nature. 2008 Apr 17;452(7189):872-6 PMID: 18421352
  18. alpha-Synuclein locus triplication causes Parkinson's disease.
    Science. 2003 Oct 31;302(5646):841 PMID: 14593171
  19. High-resolution mapping of copy-number alterations with massively parallel sequencing.
    Nat Methods. 2009 Jan;6(1):99-103 PMID: 19043412
  20. Origins and functional impact of copy number variation in the human genome.
    Nature. 2010 Apr 1;464(7289):704-12 PMID: 19812545
  21. Targeted capture and massively parallel sequencing of 12 human exomes.
    Nature. 2009 Sep 10;461(7261):272-6 PMID: 19684571
  22. Fine-scale structural variation of the human genome.
    Nat Genet. 2005 Jul;37(7):727-32 PMID: 15895083
  23. A map of human genome variation from population-scale sequencing.
    Nature. 2010 Oct 28;467(7319):1061-73 PMID: 20981092
  24. Combinatorial algorithms for structural variation detection in high-throughput sequenced genomes.
    Genome Res. 2009 Jul;19(7):1270-8 PMID: 19447966
  25. Performance comparison of exome DNA sequencing technologies.
    Nat Biotechnol. 2011 Sep 25;29(10):908-14 PMID: 21947028
  26. PEMer: a computational framework with simulation-based error models for inferring genomic structural variants from massive paired-end sequencing data.
    Genome Biol. 2009 Feb 23;10(2):R23 PMID: 19236709
  27. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
  28. Exome sequencing: the sweet spot before whole genomes.
    Hum Mol Genet. 2010 Oct 15;19(R2):R145-51 PMID: 20705737
  29. Mapping and sequencing of structural variation from eight human genomes.
    Nature. 2008 May 1;453(7191):56-64 PMID: 18451855
  30. Towards a comprehensive structural variation map of an individual human genome.
    Genome Biol. 2010;11(5):R52 PMID: 20482838
  31. Exome sequencing-based copy-number variation and loss of heterozygosity detection: ExomeCNV.
    Bioinformatics. 2011 Oct 1;27(19):2648-54 PMID: 21828086
  32. Evaluation of next generation sequencing platforms for population targeted sequencing studies.
    Genome Biol. 2009;10(3):R32 PMID: 19327155
  33. Comparative analysis of algorithms for identifying amplifications and deletions in array CGH data.
    Bioinformatics. 2005 Oct 1;21(19):3763-70 PMID: 16081473
  34. A copy number variation morbidity map of developmental delay.
    Nat Genet. 2011 Aug 14;43(9):838-46 PMID: 21841781
  35. Fast gapped-read alignment with Bowtie 2.
    Nat Methods. 2012 Mar 04;9(4):357-9 PMID: 22388286
  36. A very fast and accurate method for calling aberrations in array-CGH data.
    Biostatistics. 2010 Jul;11(3):515-8 PMID: 20207682
  37. Accurate whole human genome sequencing using reversible terminator chemistry.
    Nature. 2008 Nov 6;456(7218):53-9 PMID: 18987734
  38. Copy number variation detection and genotyping from exome sequence data.
    Genome Res. 2012 Aug;22(8):1525-32 PMID: 22585873
  39. Pindel: a pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads.
    Bioinformatics. 2009 Nov 1;25(21):2865-71 PMID: 19561018
  40. CONTRA: copy number analysis for targeted resequencing.
    Bioinformatics. 2012 May 15;28(10):1307-13 PMID: 22474122
  41. APP locus duplication causes autosomal dominant early-onset Alzheimer disease with cerebral amyloid angiopathy.
    Nat Genet. 2006 Jan;38(1):24-6 PMID: 16369530
  42. Detection of large-scale variation in the human genome.
    Nat Genet. 2004 Sep;36(9):949-51 PMID: 15286789
  43. VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing.
    Genome Res. 2012 Mar;22(3):568-76 PMID: 22300766
  44. Genome-wide loss of heterozygosity and copy number analysis in melanoma using high-density single-nucleotide polymorphism arrays.
    Cancer Res. 2007 Mar 15;67(6):2632-42 PMID: 17363583
  45. The Sequence Alignment/Map format and SAMtools.
    Bioinformatics. 2009 Aug 15;25(16):2078-9 PMID: 19505943
Article Info
Journal
Genome biology
Abbr.
Genome Biol
ISSN
1474-760X
Published
2013-00-00
Pages
R120
Language
English
Region
England
NLM ID
100960660
PMCID
PMC4053953
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]