Abstract
Although the list of completed genome sequencing projects has expanded rapidly, sequencing and analysis of expressed sequence tags (ESTs) remain a primary tool for discovery of novel genes in many eukaryotes and a key element in genome annotation. The TIGR Gene Indices (http://www.tigr.org/tdb/tgi) are a collection of 77 species-specific databases that use a highly refined protocol to analyze gene and EST sequences in an attempt to identify and characterize expressed transcripts and to present them on the Web in a user-friendly, consistent fashion. A Gene Index database is constructed for each selected organism by first clustering, then assembling EST and annotated cDNA and gene sequences from GenBank. This process produces a set of unique, high-fidelity virtual transcripts, or tentative consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to genetic and physical maps, to provide links to orthologous and paralogous genes, and as a resource for comparative and functional genomic analysis.
MeSH Terms
Animals
Base Sequence
Consensus Sequence
Databases, Genetic/trends
Eukaryotic Cells/metabolism
Expressed Sequence Tags/chemistry
Genome
Genomics
Humans
Internet
Sequence Analysis, DNA
Software
Authors & Affiliations
10 authors, click to expand affiliations / ORCID
Lee Y
The Institute for Genomic Research, 9712 Medical Center Drive, Rockville, MD 20850, USA.
[email protected]
Tsai J
Sunkara S
Karamycheva S
Pertea G
Sultana R
Antonescu V
Chan A
Cheung F
Quackenbush J
References (11)
11 references, click to expand
-
The TIGR gene indices: reconstruction and representation of expressed gene sequences.
Nucleic Acids Res. 2000 Jan 1;28(1):141-5
PMID: 10592205
-
ESTScan: a program for detecting, evaluating, and reconstructing potential coding regions in EST sequences.
Proc Int Conf Intell Syst Mol Biol. 1999;:138-48
PMID: 10786296
-
A greedy algorithm for aligning DNA sequences.
J Comput Biol. 2000 Feb-Apr;7(1-2):203-14
PMID: 10890397
-
An optimized protocol for analysis of EST sequences.
Nucleic Acids Res. 2000 Sep 15;28(18):3657-65
PMID: 10982889
-
The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species.
Nucleic Acids Res. 2001 Jan 1;29(1):159-64
PMID: 11125077
-
RESOURCERER: a database for annotating and linking microarray resources within and across species.
Genome Biol. 2001;2(11):SOFTWARE0002
PMID: 16173164
-
TIGR Gene Indices clustering tools (TGICL): a software system for fast clustering of large EST datasets.
Bioinformatics. 2003 Mar 22;19(5):651-2
PMID: 12651724
-
Selection of oligonucleotide probes for protein coding sequences.
Bioinformatics. 2003 May 1;19(7):796-802
PMID: 12724288
-
The distributed annotation system.
BMC Bioinformatics. 2001;2:7
PMID: 11667947
-
CAP3: A DNA sequence assembly program.
Genome Res. 1999 Sep;9(9):868-77
PMID: 10508846
-
Cross-referencing eukaryotic genomes: TIGR Orthologous Gene Alignments (TOGA).
Genome Res. 2002 Mar;12(3):493-502
PMID: 11875039