Home LiteratureArticle Details
PMID: 18976482 Published · epublish English Journal Article Research Support, U.S. Gov't, Non-P.H.S.

A new method to compute K-mer frequencies and its application to annotate large repetitive plant genomes.

BMC genomics ·Vol. 9 ·2008-10-31 ·Pages 517

Kurtz S, Narechania A, Stein JC, Ware D

Abstract

The challenges of accurate gene prediction and enumeration are further aggravated in large genomes that contain highly repetitive transposable elements (TEs). Yet TEs play a substantial role in genome evolution and are themselves an important subject of study. Repeat annotation, based on counting occurrences of k-mers, has been previously used to distinguish TEs from low-copy genic regions; but currently available software solutions are impractical due to high memory requirements or specialization for specific user-tasks. Here we introduce the Tallymer software, a flexible and memory-efficient collection of programs for k-mer counting and indexing of large sequence sets. Unlike previous methods, Tallymer is based on enhanced suffix arrays. This gives a much larger flexibility concerning the choice of the k-mer size. Tallymer can process large data sizes of several billion bases. We used it in a variety of applications to study the genomes of maize and other plant species. In particular, Tallymer was used to index a set of whole genome shotgun sequences from maize (B73) (total size 109 bp.). We analyzed k-mer frequencies for a wide range of k. At this low genome coverage ( approximately 0.45x) highly repetitive 20-mers constituted 44% of the genome but represented only 1% of all possible k-mers. Similar low-complexity was seen in the repeat fractions of sorghum and rice. When applying our method to other maize data sets, High-C0t derived sequences showed the greatest enrichment for low-copy sequences. Among annotated TEs, the most highly repetitive were of the Ty3/gypsy class of retrotransposons, followed by the Ty1/copia class, and DNA transposons. Among expressed sequence tags (EST), a notable fraction contained high-copy k-mers, suggesting that transposons are still active in maize. Retrotransposons in Mo17 and McC cultivars were readily detected using the B73 20-mer frequency index, indicating their conservation despite extensive rearrangement across cultivars. Among one hundred annotated bacterial artificial chromosomes (BACs), k-mer frequency could be used to detect transposon-encoded genes with 92% sensitivity, compared to 96% using alignment-based repeat masking, while both methods showed 92% specificity. The Tallymer software was effective in a variety of applications to aid genome annotation in maize, despite limitations imposed by the relatively low coverage of sequence available. For more information on the software, see http://www.zbh.uni-hamburg.de/Tallymer.

MeSH Terms
Computational Biology/methods DNA Transposable Elements Genome, Plant Genomics/methods Methods Oryza Software Sorghum Zea mays
Chemicals
DNA Transposable Elements
Authors & Affiliations
4 authors, click to expand affiliations / ORCID
Kurtz Stefan
Center for Bioinformatics, University of Hamburg, Bundesstrasse 43, 20146 Hamburg, Germany. [email protected]
Narechania Apurva
Stein Joshua C
Ware Doreen
References (38)
38 references, click to expand
  1. Structure and architecture of the maize genome.
    Plant Physiol. 2005 Dec;139(4):1612-24 PMID: 16339807
  2. Whole-genome validation of high-information-content fingerprinting.
    Plant Physiol. 2005 Sep;139(1):27-38 PMID: 16166258
  3. ReAS: Recovery of ancestral sequences for transposable elements from the unassembled reads of a whole genome shotgun.
    PLoS Comput Biol. 2005 Sep;1(4):e43 PMID: 16184192
  4. The genome of black cottonwood, Populus trichocarpa (Torr. & Gray).
    Science. 2006 Sep 15;313(5793):1596-604 PMID: 16973872
  5. Whole-genome re-sequencing.
    Curr Opin Genet Dev. 2006 Dec;16(6):545-52 PMID: 17055251
  6. Striking similarities in the genomic distribution of tandemly arrayed genes in Arabidopsis and rice.
    PLoS Comput Biol. 2006 Sep 1;2(9):e115 PMID: 16948529
  7. MIPSPlantsDB--plant database resource for integrative and comparative plant genome research.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D834-40 PMID: 17202173
  8. The grapevine genome sequence suggests ancestral hexaploidization in major angiosperm phyla.
    Nature. 2007 Sep 27;449(7161):463-7 PMID: 17721507
  9. The impact of next-generation sequencing technology on genetics.
    Trends Genet. 2008 Mar;24(3):133-41 PMID: 18262675
  10. Low-pass shotgun sequencing of the barley genome facilitates rapid identification of genes, conserved non-coding sequences and novel repeats.
    BMC Genomics. 2008;9:518 PMID: 18976483
  11. Analysis of the genome sequence of the flowering plant Arabidopsis thaliana.
    Nature. 2000 Dec 14;408(6814):796-815 PMID: 11130711
  12. Design of a compartmentalized shotgun assembler for the human genome.
    Bioinformatics. 2001;17 Suppl 1:S132-9 PMID: 11473002
  13. A clustering method for repeat analysis in DNA sequences.
    Genome Biol. 2001;2(8):RESEARCH0027 PMID: 11532211
  14. REPuter: the manifold applications of repeat analysis on a genomic scale.
    Nucleic Acids Res. 2001 Nov 15;29(22):4633-42 PMID: 11713313
  15. Access to the maize genome: an integrated physical and genetic map.
    Plant Physiol. 2002 Jan;128(1):9-12 PMID: 11788746
  16. A draft sequence of the rice genome (Oryza sativa L. ssp. indica).
    Science. 2002 Apr 5;296(5565):79-92 PMID: 11935017
  17. A draft sequence of the rice genome (Oryza sativa L. ssp. japonica).
    Science. 2002 Apr 5;296(5565):92-100 PMID: 11935018
  18. Transposable elements, genes and recombination in a 215-kb contig from wheat chromosome 5A(m).
    Funct Integr Genomics. 2002 May;2(1-2):70-80 PMID: 12021852
  19. Automated de novo identification of repeat sequence families in sequenced genomes.
    Genome Res. 2002 Aug;12(8):1269-76 PMID: 12176934
  20. FORRepeats: detects repeats on entire chromosomes and between genomes.
    Bioinformatics. 2003 Feb 12;19(3):319-26 PMID: 12584116
  21. Annotating large genomes with exact word matches.
    Genome Res. 2003 Oct;13(10):2306-15 PMID: 12975312
  22. Structure and evolution of the Cinful retrotransposon family of maize.
    Genome. 2003 Oct;46(5):745-52 PMID: 14608391
  23. PlantGDB, plant genome database and analysis tools.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D354-9 PMID: 14681433
  24. Enrichment of gene-coding sequences in maize by genome filtration.
    Science. 2003 Dec 19;302(5653):2118-20 PMID: 14684821
  25. Comparative analysis of a Brassica BAC clone containing several major aliphatic glucosinolate genes with its corresponding Arabidopsis sequence.
    Genome. 2004 Aug;47(4):666-79 PMID: 15284871
  26. De novo repeat classification and fragment assembly.
    Genome Res. 2004 Sep;14(9):1786-96 PMID: 15342561
  27. Selfish genes, the phenotype paradigm and genome evolution.
    Nature. 1980 Apr 17;284(5757):601-3 PMID: 6245369
  28. Selfish DNA: the ultimate parasite.
    Nature. 1980 Apr 17;284(5757):604-7 PMID: 7366731
  29. A method of comparing the areas under receiver operating characteristic curves derived from the same cases.
    Radiology. 1983 Sep;148(3):839-43 PMID: 6878708
  30. The paleontology of intergene retrotransposons of maize.
    Nat Genet. 1998 Sep;20(1):43-5 PMID: 9731528
  31. Genome evolution in the genus Sorghum (Poaceae).
    Ann Bot. 2005 Jan;95(1):219-27 PMID: 15596469
  32. RAP: a new computer program for de novo identification of repeated sequences in whole genomes.
    Bioinformatics. 2005 Mar 1;21(5):582-8 PMID: 15374857
  33. PILER: identification and classification of genomic repeats.
    Bioinformatics. 2005 Jun;21 Suppl 1:i152-8 PMID: 15961452
  34. De novo identification of repeat families in large genomes.
    Bioinformatics. 2005 Jun;21 Suppl 1:i351-8 PMID: 15961478
  35. Gene movement by Helitron transposons contributes to the haplotype variability of maize.
    Proc Natl Acad Sci U S A. 2005 Jun 21;102(25):9068-73 PMID: 15951422
  36. The map-based sequence of the rice genome.
    Nature. 2005 Aug 11;436(7052):793-800 PMID: 16100779
  37. Genome sequencing in microfabricated high-density picolitre reactors.
    Nature. 2005 Sep 15;437(7057):376-80 PMID: 16056220
  38. The TIGR Maize Database.
    Nucleic Acids Res. 2006 Jan 1;34(Database issue):D771-6 PMID: 16381977
Article Info
Journal
BMC genomics
Abbr.
BMC Genomics
ISSN
1471-2164
Published
2008-10-31
Epub
2008-00-31
Pages
517
Language
English
Region
England
NLM ID
100965258
PMCID
PMC2613927
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]