Home LiteratureArticle Details
PMID: 23299411 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Rice Annotation Project Database (RAP-DB): an integrative and interactive database for rice genomics.

Plant & cell physiology ·Vol. 54 ·No. 2 ·2013-02-00 ·Pages e6

Sakai H, Lee SS, Tanaka T, Numa H, Kim J, Kawahara Y, Wakimoto H, Yang CC, Iwamoto M, Abe T, Yamada Y, Muto A, Inokuchi H, Ikemura T, Matsumoto T, Sasaki T, Itoh T

Abstract

The Rice Annotation Project Database (RAP-DB, http://rapdb.dna.affrc.go.jp/) has been providing a comprehensive set of gene annotations for the genome sequence of rice, Oryza sativa (japonica group) cv. Nipponbare. Since the first release in 2005, RAP-DB has been updated several times along with the genome assembly updates. Here, we present our newest RAP-DB based on the latest genome assembly, Os-Nipponbare-Reference-IRGSP-1.0 (IRGSP-1.0), which was released in 2011. We detected 37,869 loci by mapping transcript and protein sequences of 150 monocot species. To provide plant researchers with highly reliable and up to date rice gene annotations, we have been incorporating literature-based manually curated data, and 1,626 loci currently incorporate literature-based annotation data, including commonly used gene names or gene symbols. Transcriptional activities are shown at the nucleotide level by mapping RNA-Seq reads derived from 27 samples. We also mapped the Illumina reads of a Japanese leading japonica cultivar, Koshihikari, and a Chinese indica cultivar, Guangluai-4, to the genome and show alignments together with the single nucleotide polymorphisms (SNPs) and gene functional annotations through a newly developed browser, Short-Read Assembly Browser (S-RAB). We have developed two satellite databases, Plant Gene Family Database (PGFD) and Integrative Database of Cereal Gene Phylogeny (IDCGP), which display gene family and homologous gene relationships among diverse plant species. RAP-DB and the satellite databases offer simple and user-friendly web interfaces, enabling plant and genome researchers to access the data easily and facilitating a broad range of plant research topics.

MeSH Terms
Base Sequence Databases, Genetic Gene Expression Profiling Genes, Plant Genetic Loci Genomics/methods Microsatellite Repeats Molecular Sequence Annotation Molecular Sequence Data Oryza/classification,genetics Phylogeny Polymorphism, Single Nucleotide Search Engine Sequence Homology
Authors & Affiliations
17 authors, click to expand affiliations / ORCID
Sakai Hiroaki
Agrogenomics Research Center, National Institute of Agrobiological Sciences, Tsukuba, Ibaraki, Japan.
Lee Sung Shin
Tanaka Tsuyoshi
Numa Hisataka
Kim Jungsok
Kawahara Yoshihiro
Wakimoto Hironobu
Yang Ching-chia
Iwamoto Masao
Abe Takashi
Yamada Yuko
Muto Akira
Inokuchi Hachiro
Ikemura Toshimichi
Matsumoto Takashi
Sasaki Takuji
Itoh Takeshi
References (54)
54 references, click to expand
  1. Genomics and bioinformatics resources for crop improvement.
    Plant Cell Physiol. 2010 Apr;51(4):497-523 PMID: 20208064
  2. Fast and accurate short read alignment with Burrows-Wheeler transform.
    Bioinformatics. 2009 Jul 15;25(14):1754-60 PMID: 19451168
  3. MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform.
    Nucleic Acids Res. 2002 Jul 15;30(14):3059-66 PMID: 12136088
  4. Computational and experimental analysis of microsatellites in rice (Oryza sativa L.): frequency, length variation, transposon associations, and genetic marker potential.
    Genome Res. 2001 Aug;11(8):1441-52 PMID: 11483586
  5. Oryzabase. An integrated biological and genome information database for rice.
    Plant Physiol. 2006 Jan;140(1):12-7 PMID: 16403737
  6. Development of 5006 full-length CDNAs in barley: a tool for accessing cereal genomics resources.
    DNA Res. 2009 Apr;16(2):81-9 PMID: 19150987
  7. The neighbor-joining method: a new method for reconstructing phylogenetic trees.
    Mol Biol Evol. 1987 Jul;4(4):406-25 PMID: 3447015
  8. NCBI Reference Sequences (RefSeq): current status, new features and genome annotation policy.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D130-5 PMID: 22121212
  9. ATTED-II updates: condition-specific gene coexpression to extend coexpression analyses and applications to a broad range of flowering plants.
    Plant Cell Physiol. 2011 Feb;52(2):213-9 PMID: 21217125
  10. A map of rice genome variation reveals the origin of cultivated rice.
    Nature. 2012 Oct 25;490(7421):497-501 PMID: 23034647
  11. BLAT--the BLAST-like alignment tool.
    Genome Res. 2002 Apr;12(4):656-64 PMID: 11932250
  12. Efficient decoding algorithms for generalized hidden Markov model gene finders.
    BMC Bioinformatics. 2005 Jan 24;6:16 PMID: 15667658
  13. TriFLDB: a database of clustered full-length coding sequences from Triticeae with applications to comparative grass genomics.
    Plant Physiol. 2009 Jul;150(3):1135-46 PMID: 19448038
  14. The map-based sequence of the rice genome.
    Nature. 2005 Aug 11;436(7052):793-800 PMID: 16100779
  15. Clustal W and Clustal X version 2.0.
    Bioinformatics. 2007 Nov 1;23(21):2947-8 PMID: 17846036
  16. The TIGR Plant Repeat Databases: a collective resource for the identification of repetitive sequences in plants.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D360-3 PMID: 14681434
  17. Curated genome annotation of Oryza sativa ssp. japonica and comparative genome analysis with Arabidopsis thaliana.
    Genome Res. 2007 Feb;17(2):175-83 PMID: 17210932
  18. Massive gene losses in Asian cultivated rice unveiled by comparative genome analysis.
    BMC Genomics. 2010 Feb 19;11:121 PMID: 20167122
  19. An SNP caused loss of seed shattering during rice domestication.
    Science. 2006 Jun 2;312(5778):1392-6 PMID: 16614172
  20. JIGSAW: integration of multiple sources of evidence for gene prediction.
    Bioinformatics. 2005 Sep 15;21(18):3596-603 PMID: 16076884
  21. The TIGR Rice Genome Annotation Resource: improvements and new features.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D883-7 PMID: 17145706
  22. RiceXPro: a platform for monitoring gene expression in japonica rice grown under natural field conditions.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D1141-8 PMID: 21045061
  23. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data.
    Genome Res. 2010 Sep;20(9):1297-303 PMID: 20644199
  24. Efficient plant gene identification based on interspecies mapping of full-length cDNAs.
    DNA Res. 2010 Oct;17(5):271-9 PMID: 20668003
  25. DDBJ launches a new archive database with analytical tools for next-generation sequence data.
    Nucleic Acids Res. 2010 Jan;38(Database issue):D33-8 PMID: 19850725
  26. Genome-wide association studies of 14 agronomic traits in rice landraces.
    Nat Genet. 2010 Nov;42(11):961-7 PMID: 20972439
  27. Comprehensive sequence analysis of 24,783 barley full-length cDNAs derived from 12 clone libraries.
    Plant Physiol. 2011 May;156(1):20-8 PMID: 21415278
  28. KEGG for integration and interpretation of large-scale molecular data sets.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D109-14 PMID: 22080510
  29. Insights into corn genes derived from large-scale cDNA sequencing.
    Plant Mol Biol. 2009 Jan;69(1-2):179-94 PMID: 18937034
  30. EST_GENOME: a program to align spliced DNA sequences to unspliced genomic DNA.
    Comput Appl Biosci. 1997 Aug;13(4):477-8 PMID: 9283765
  31. tRNADB-CE: tRNA gene database curated manually by experts.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D163-8 PMID: 18842632
  32. Reorganizing the protein space at the Universal Protein Resource (UniProt).
    Nucleic Acids Res. 2012 Jan;40(Database issue):D71-5 PMID: 22102590
  33. TopHat: discovering splice junctions with RNA-Seq.
    Bioinformatics. 2009 May 1;25(9):1105-11 PMID: 19289445
  34. OryzaExpress: an integrated database of gene expression networks and omics annotations in rice.
    Plant Cell Physiol. 2011 Feb;52(2):220-9 PMID: 21186175
  35. Highly diversified molecular evolution of downstream transcription start sites in rice and Arabidopsis.
    Plant Physiol. 2009 Mar;149(3):1316-24 PMID: 19118127
  36. Fine definition of the pedigree haplotypes of closely related rice cultivars by means of genome-wide discovery of single-nucleotide polymorphisms.
    BMC Genomics. 2010 Apr 27;11:267 PMID: 20423466
  37. GeneMark.hmm: new solutions for gene finding.
    Nucleic Acids Res. 1998 Feb 15;26(4):1107-15 PMID: 9461475
  38. The Rice Annotation Project Database (RAP-DB): 2008 update.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D1028-33 PMID: 18089549
  39. Annotation, submission and screening of repetitive elements in Repbase: RepbaseSubmitter and Censor.
    BMC Bioinformatics. 2006 Oct 25;7:474 PMID: 17064419
  40. Collection and comparative analysis of 1888 full-length cDNAs from wild rice Oryza rufipogon Griff. W1943.
    DNA Res. 2008 Oct;15(5):285-95 PMID: 18687674
  41. Retrogenes in rice (Oryza sativa L. ssp. japonica) exhibit correlated expression with their source genes.
    Genome Biol Evol. 2011;3:1357-68 PMID: 22042334
  42. Phytozome: a comparative platform for green plant genomics.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D1178-86 PMID: 22110026
  43. Massive parallel sequencing of mRNA in identification of unannotated salinity stress-inducible transcripts in rice (Oryza sativa L.).
    BMC Genomics. 2010 Dec 02;11:683 PMID: 21122150
  44. Loss of function of a proline-containing protein confers durable disease resistance in rice.
    Science. 2009 Aug 21;325(5943):998-1001 PMID: 19696351
  45. Sequencing, mapping, and analysis of 27,455 maize full-length cDNAs.
    PLoS Genet. 2009 Nov;5(11):e1000740 PMID: 19936069
  46. The Arabidopsis Information Resource (TAIR): improved gene annotation and new tools.
    Nucleic Acids Res. 2012 Jan;40(Database issue):D1202-10 PMID: 22140109
  47. Sequence mapping by electronic PCR
    Genome Res. 1997 May;7(5):541-50 PMID: 9149949
  48. miRBase: integrating microRNA annotation and deep-sequencing data.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D152-7 PMID: 21037258
  49. TigrScan and GlimmerHMM: two open source ab initio eukaryotic gene-finders.
    Bioinformatics. 2004 Nov 1;20(16):2878-9 PMID: 15145805
  50. Gramene database in 2010: updates and extensions.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D1085-94 PMID: 21076153
  51. Annotation and expression profile analysis of 2073 full-length cDNAs from stress-induced maize (Zea mays L.) seedlings.
    Plant J. 2006 Dec;48(5):710-27 PMID: 17076806
  52. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  53. MIPSPlantsDB--plant database resource for integrative and comparative plant genome research.
    Nucleic Acids Res. 2007 Jan;35(Database issue):D834-40 PMID: 17202173
  54. Genome-wide association study of flowering time and grain yield traits in a worldwide collection of rice germplasm.
    Nat Genet. 2011 Dec 04;44(1):32-9 PMID: 22138690
Article Info
Journal
Plant & cell physiology
Abbr.
Plant Cell Physiol
ISSN
1471-9053
Published
2013-02-00
Epub
2013-00-07
Pages
e6
Language
English
Region
Japan
NLM ID
9430925
PMCID
PMC3583025
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]