Home LiteratureArticle Details
PMID: 22323520 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, U.S. Gov't, Non-P.H.S.

Summarizing and correcting the GC content bias in high-throughput sequencing.

Nucleic acids research ·Vol. 40 ·No. 10 ·2012-05-00 ·Pages e72

Benjamini Y, Speed TP

Abstract

GC content bias describes the dependence between fragment count (read coverage) and GC content found in Illumina sequencing data. This bias can dominate the signal of interest for analyses that focus on measuring fragment abundance within a genome, such as copy number estimation (DNA-seq). The bias is not consistent between samples; and there is no consensus as to the best methods to remove it in a single sample. We analyze regularities in the GC bias patterns, and find a compact description for this unimodal curve family. It is the GC content of the full DNA fragment, not only the sequenced read, that most influences fragment count. This GC effect is unimodal: both GC-rich fragments and AT-rich fragments are underrepresented in the sequencing results. This empirical evidence strengthens the hypothesis that PCR is the most important cause of the GC bias. We propose a model that produces predictions at the base pair level, allowing strand-specific GC-effect correction regardless of the downstream smoothing or binning. These GC modeling considerations can inform other high-throughput sequencing analyses such as ChIP-seq and RNA-seq.

MeSH Terms
Base Composition DNA/chemistry Genome, Human High-Throughput Nucleotide Sequencing/methods Humans Models, Genetic Poisson Distribution Sequence Analysis, DNA/methods
Chemicals
DNA
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Benjamini Yuval
Department of Statistics, University of California, Berkeley, CA 94720, USA. [email protected]
Speed Terence P
References (19)
19 references, click to expand
  1. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  2. High-resolution mapping of copy-number alterations with massively parallel sequencing.
    Nat Methods. 2009 Jan;6(1):99-103 PMID: 19043412
  3. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome.
    Genome Biol. 2009;10(3):R25 PMID: 19261174
  4. Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
    Nucleic Acids Res. 2008 Sep;36(16):e105 PMID: 18660515
  5. A large genome center's improvements to the Illumina sequencing system.
    Nat Methods. 2008 Dec;5(12):1005-10 PMID: 19034268
  6. Sensitive and accurate detection of copy number variants using read depth of coverage.
    Genome Res. 2009 Sep;19(9):1586-92 PMID: 19657104
  7. A Statistical Framework for the Analysis of ChIP-Seq Data.
    J Am Stat Assoc. 2011;106(495):891-903 PMID: 26478641
  8. Control-free calling of copy number alterations in deep-sequencing data using GC-content normalization.
    Bioinformatics. 2011 Jan 15;27(2):268-9 PMID: 21081509
  9. Impact of chromatin structures on DNA processing for genomic analyses.
    PLoS One. 2009 Aug 20;4(8):e6700 PMID: 19693276
  10. CNAseg--a novel framework for identification of copy number changes in cancer from second-generation sequencing data.
    Bioinformatics. 2010 Dec 15;26(24):3051-8 PMID: 20966003
  11. Mapping short DNA sequencing reads and calling variants using mapping quality scores.
    Genome Res. 2008 Nov;18(11):1851-8 PMID: 18714091
  12. ReadDepth: a parallel R package for detecting copy number alterations from short sequencing reads.
    PLoS One. 2011 Jan 31;6(1):e16327 PMID: 21305028
  13. Model-based quality assessment and base-calling for second-generation sequencing data.
    Biometrics. 2010 Sep;66(3):665-74 PMID: 19912177
  14. Accurate whole human genome sequencing using reversible terminator chemistry.
    Nature. 2008 Nov 6;456(7218):53-9 PMID: 18987734
  15. Biases in Illumina transcriptome sequencing caused by random hexamer priming.
    Nucleic Acids Res. 2010 Jul;38(12):e131 PMID: 20395217
  16. Amplification-free Illumina sequencing-library preparation facilitates improved mapping and assembly of (G+C)-biased genomes.
    Nat Methods. 2009 Apr;6(4):291-5 PMID: 19287394
  17. Analyzing and minimizing PCR amplification bias in Illumina sequencing libraries.
    Genome Biol. 2011;12(2):R18 PMID: 21338519
  18. Systematic bias in high-throughput sequencing data and its correction by BEADS.
    Nucleic Acids Res. 2011 Aug;39(15):e103 PMID: 21646344
  19. Sequence-specific error profile of Illumina sequencers.
    Nucleic Acids Res. 2011 Jul;39(13):e90 PMID: 21576222
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2012-05-00
Epub
2012-00-09
Pages
e72
Language
English
Region
England
NLM ID
0411011
PMCID
PMC3378858
Subset
IM
Grants
NCI NIH HHS · 3U24CA143799-02S1 · United States
NIGMS NIH HHS · 5R01 GM083084-03 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]