Abstract
NCBI's reference sequence (RefSeq) database (http://www.ncbi.nlm.nih.gov/RefSeq/) is a curated non-redundant collection of sequences representing genomes, transcripts and proteins. The database includes 3774 organisms spanning prokaryotes, eukaryotes and viruses, and has records for 2,879,860 proteins (RefSeq release 19). RefSeq records integrate information from multiple sources, when additional data are available from those sources and therefore represent a current description of the sequence and its features. Annotations include coding regions, conserved domains, tRNAs, sequence tagged sites (STS), variation, references, gene and protein product names, and database cross-references. Sequence is reviewed and features are added using a combined approach of collaboration and other input from the scientific community, prediction, propagation from GenBank and curation by NCBI staff. The format of all RefSeq records is validated, and an increasing number of tests are being applied to evaluate the quality of sequence and annotation, especially in the context of complete genomic sequence.
MeSH Terms
Amino Acid Sequence
Base Sequence
Databases, Nucleic Acid
Databases, Protein
Genome
Internet
National Library of Medicine (U.S.)
Quality Control
RNA, Messenger/chemistry
Reference Standards
Sequence Analysis, DNA/standards
Sequence Analysis, Protein/standards
Sequence Analysis, RNA/standards
United States
User-Computer Interface
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Pruitt Kim D
National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health, Rm 6An.12J, 45 Center Drive, Bethesda, MD 20892-6510, USA.
[email protected]
Tatusova Tatiana
Maglott Donna R
References (12)
12 references, click to expand
-
Complete genomes in WWW Entrez: data representation and analysis.
Bioinformatics. 1999 Jul-Aug;15(7-8):536-43
PMID: 10487861
-
FlyBase: genes and gene models.
Nucleic Acids Res. 2005 Jan 1;33(Database issue):D390-5
PMID: 15608223
-
WormBase: better software, richer content.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D475-8
PMID: 16381915
-
The Mouse Genome Database (MGD): updates and enhancements.
Nucleic Acids Res. 2006 Jan 1;34(Database issue):D562-7
PMID: 16381933
-
Entrez Gene: gene-centered information at NCBI.
Nucleic Acids Res. 2007 Jan;35(Database issue):D26-31
PMID: 17148475
-
GenBank.
Nucleic Acids Res. 2007 Jan;35(Database issue):D21-5
PMID: 17202161
-
Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
Nucleic Acids Res. 1997 Sep 1;25(17):3389-402
PMID: 9254694
-
The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community.
Nucleic Acids Res. 2003 Jan 1;31(1):224-8
PMID: 12519987
-
Generation of protein isoform diversity by alternative initiation of translation at non-AUG codons.
Biol Cell. 2003 May-Jun;95(3-4):169-78
PMID: 12867081
-
Regulation of gene expression by stop codon recoding: selenocysteine.
Gene. 2003 Jul 17;312:17-25
PMID: 12909337
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
Entrez: molecular biology database and retrieval system.
Methods Enzymol. 1996;266:141-62
PMID: 8743683