Abstract
Numerous types of DNA variation exist, ranging from SNPs to larger structural alterations such as copy number variants (CNVs) and inversions. Alignment of DNA sequence from different sources has been used to identify SNPs and intermediate-sized variants (ISVs). However, only a small proportion of total heterogeneity is characterized, and little is known of the characteristics of most smaller-sized (<50 kb) variants. Here we show that genome assembly comparison is a robust approach for identification of all classes of genetic variation. Through comparison of two human assemblies (Celera's R27c compilation and the Build 35 reference sequence), we identified megabases of sequence (in the form of 13,534 putative non-SNP events) that were absent, inverted or polymorphic in one assembly. Database comparison and laboratory experimentation further demonstrated overlap or validation for 240 variable regions and confirmed >1.5 million SNPs. Some differences were simple insertions and deletions, but in regions containing CNVs, segmental duplication and repetitive DNA, they were more complex. Our results uncover substantial undescribed variation in humans, highlighting the need for comprehensive annotation strategies to fully interpret genome scanning and personalized sequencing projects.
MeSH Terms
Base Sequence
DNA/genetics
Genetic Variation
Genome, Human
Genomics
Humans
In Situ Hybridization, Fluorescence
Polymerase Chain Reaction
Sequence Alignment
Authors & Affiliations
20 authors, click to expand affiliations / ORCID
Khaja Razi
Program in Genetics and Genomic Biology, The Hospital for Sick Children and Department of Molecular and Medical Genetics, University of Toronto and The Centre for Applied Genomics, MaRS Centre, Toronto, Ontario, M5G 1L7, Canada.
Zhang Junjun
MacDonald Jeffrey R
He Yongshu
Joseph-George Ann M
Wei John
Rafiq Muhammad A
Qian Cheng
Shago Mary
Pantano Lorena
Aburatani Hiroyuki
Jones Keith
Redon Richard
Hurles Matthew
Armengol Lluis
Estivill Xavier
Mural Richard J
Lee Charles
Scherer Stephen W
Feuk Lars
References (29)
29 references, click to expand
-
Single nucleotide polymorphisms (SNPs) that map to gaps in the human SNP map.
Nucleic Acids Res. 2003 Aug 15;31(16):4910-6
PMID: 12907734
-
BLAT--the BLAST-like alignment tool.
Genome Res. 2002 Apr;12(4):656-64
PMID: 11932250
-
Structural variation in the human genome.
Nat Rev Genet. 2006 Feb;7(2):85-97
PMID: 16418744
-
Discovery of human inversion polymorphisms by comparative analysis of human and chimpanzee DNA sequence assemblies.
PLoS Genet. 2005 Oct;1(4):e56
PMID: 16254605
-
A 1.5 million-base pair inversion polymorphism in families with Williams-Beuren syndrome.
Nat Genet. 2001 Nov;29(3):321-5
PMID: 11685205
-
On the sequencing and assembly of the human genome.
Proc Natl Acad Sci U S A. 2002 Apr 2;99(7):4145-6
PMID: 11904395
-
Recent segmental duplications in the human genome.
Science. 2002 Aug 9;297(5583):1003-7
PMID: 12169732
-
Toward the 1,000 dollars human genome.
Pharmacogenomics. 2005 Jun;6(4):373-82
PMID: 16004555
-
A greedy algorithm for aligning DNA sequences.
J Comput Biol. 2000 Feb-Apr;7(1-2):203-14
PMID: 10890397
-
The human genome browser at UCSC.
Genome Res. 2002 Jun;12(6):996-1006
PMID: 12045153
-
dbRIP: a highly integrated database of retrotransposon insertion polymorphisms in humans.
Hum Mutat. 2006 Apr;27(4):323-9
PMID: 16511833
-
Fine-scale structural variation of the human genome.
Nat Genet. 2005 Jul;37(7):727-32
PMID: 15895083
-
Global variation in copy number in the human genome.
Nature. 2006 Nov 23;444(7118):444-54
PMID: 17122850
-
The DNA sequence and comparative analysis of human chromosome 5.
Nature. 2004 Sep 16;431(7006):268-74
PMID: 15372022
-
On the sequencing of the human genome.
Proc Natl Acad Sci U S A. 2002 Mar 19;99(6):3712-6
PMID: 11880605
-
Advanced sequencing technologies: methods and goals.
Nat Rev Genet. 2004 May;5(5):335-44
PMID: 15143316
-
More on the sequencing of the human genome.
Proc Natl Acad Sci U S A. 2003 Mar 18;100(6):3022-4; author reply 3025-6
PMID: 12631699
-
A new mathematical model for relative quantification in real-time RT-PCR.
Nucleic Acids Res. 2001 May 1;29(9):e45
PMID: 11328886
-
Detection of large-scale variation in the human genome.
Nat Genet. 2004 Sep;36(9):949-51
PMID: 15286789
-
NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4
PMID: 15608248
-
A general approach to single-nucleotide polymorphism discovery.
Nat Genet. 1999 Dec;23(4):452-6
PMID: 10581034
-
The DNA sequence of human chromosome 7.
Nature. 2003 Jul 10;424(6945):157-64
PMID: 12853948
-
Gene sequencing. The race for the $1000 genome.
Science. 2006 Mar 17;311(5767):1544-6
PMID: 16543431
-
Whole-genome shotgun assembly and comparison of human genome assemblies.
Proc Natl Acad Sci U S A. 2004 Feb 17;101(7):1916-21
PMID: 14769938
-
Genome-wide detection of segmental duplications and potential assembly errors in the human genome sequence.
Genome Biol. 2003;4(4):R25
PMID: 12702206
-
Human chromosome 7: DNA sequence and biology.
Science. 2003 May 2;300(5620):767-72
PMID: 12690205
-
The sequence of the human genome.
Science. 2001 Feb 16;291(5507):1304-51
PMID: 11181995
-
Initial sequencing and analysis of the human genome.
Nature. 2001 Feb 15;409(6822):860-921
PMID: 11237011
-
The independence of our genome assemblies.
Proc Natl Acad Sci U S A. 2003 Mar 18;100(6):3025-6
PMID: 16576752