Abstract
The recent release of the first draft of the human genome provides an unprecedented opportunity to integrate human genes and their functions in a complete positional context. However, at least three significant technical hurdles remain: first, to assemble a complete and nonredundant human transcript index; second, to accurately place the individual transcript indices on the human genome; and third, to functionally annotate all human genes. Here, we report the extension of the UNIGENE database through the assembly of its sequence clusters into nonredundant sequence contigs. Each resulting consensus was aligned to the human genome draft. A unique location for each transcript within the human genome was determined by the integration of the restriction fingerprint, assembled genomic contig, and radiation hybrid (RH) maps. A total of 59,500 UNIGENE clusters were mapped on the basis of at least three independent criteria as compared with the 30,000 human genes/ESTs currently mapped in Genemap'99. Finally, the extension of the human transcript consensus in this study enabled a greater number of putative functional assignments than the 11,000 annotated entries in UNIGENE. This study reports a draft physical map with annotations for a majority of the human transcripts, called the Human Index of Nonredundant Transcripts (HINT). Such information can be immediately applied to the discovery of new genes and the identification of candidate genes for positional cloning.
MeSH Terms
Alleles
Alternative Splicing/genetics
Computational Biology/methods
Consensus Sequence/genetics
Databases, Factual
Genes/genetics
Genome, Human
Human Genome Project
Humans
Multigene Family/genetics
Sequence Alignment/methods
Authors & Affiliations
17 authors, click to expand affiliations / ORCID
Zhuo D
Bioinformatics Group, James Cancer Hospital and Solove Research Institute, The Ohio State University, Columbus, Ohio 43210, USA.
Zhao W D
Wright F A
Yang H Y
Wang J P
Sears R
Baer T
Kwon D H
Gordon D
Gibbs S
Dai D
Yang Q
Spitzner J
Krahe R
Stredney D
Stutz A
Yuan B
References (31)
31 references, click to expand
-
A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base.
Genome Res. 1999 Nov;9(11):1143-55
PMID: 10568754
-
Reliable identification of large numbers of candidate SNPs from public EST data.
Nat Genet. 1999 Mar;21(3):323-5
PMID: 10080189
-
The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
Nucleic Acids Res. 2000 Jan 1;28(1):45-8
PMID: 10592178
-
The TIGR gene indices: reconstruction and representation of expressed gene sequences.
Nucleic Acids Res. 2000 Jan 1;28(1):141-5
PMID: 10592205
-
The Pfam protein families database.
Nucleic Acids Res. 2000 Jan 1;28(1):263-6
PMID: 10592242
-
Frequent alternative splicing of human genes.
Genome Res. 1999 Dec;9(12):1288-93
PMID: 10613851
-
Representation of functional information in the SWISS-PROT data bank.
Bioinformatics. 1999 Dec;15(12):1066-7
PMID: 10746001
-
Analysis of expressed sequence tags indicates 35,000 human genes.
Nat Genet. 2000 Jun;25(2):232-4
PMID: 10835644
-
Gene index analysis of the human genome estimates approximately 120,000 genes.
Nat Genet. 2000 Jun;25(2):239-40
PMID: 10835646
-
Repeat polymorphisms within gene regions: phenotypic and evolutionary implications.
Am J Hum Genet. 2000 Aug;67(2):345-56
PMID: 10889045
-
Genome-wide analysis of single-nucleotide polymorphisms in human expressed sequences.
Nat Genet. 2000 Oct;26(2):233-6
PMID: 11017085
-
Basic local alignment search tool.
J Mol Biol. 1990 Oct 5;215(3):403-10
PMID: 2231712
-
Two acetyl-CoA acetyltransferase genes located in the t-complex region of mouse chromosome 17 partially overlap the Tcp-1 and Tcp-1x genes.
Genomics. 1993 Nov;18(2):195-8
PMID: 7904580
-
Gene discovery in dbEST.
Science. 1994 Sep 30;265(5181):1993-4
PMID: 8091218
-
False association of human ESTs.
Nat Genet. 1994 Dec;8(4):321-2
PMID: 7894479
-
ESTablishing a human transcript map.
Nat Genet. 1995 Aug;10(4):369-71
PMID: 7670480
-
Initial assessment of human gene diversity and expression patterns based upon 83 million nucleotides of cDNA sequence.
Nature. 1995 Sep 28;377(6547 Suppl):3-174
PMID: 7566098
-
The Genexpress Index: a resource for gene discovery and the genic map of the human genome.
Genome Res. 1995 Oct;5(3):272-304
PMID: 8593614
-
Generation and analysis of 280,000 human expressed sequence tags.
Genome Res. 1996 Sep;6(9):807-28
PMID: 8889549
-
Toward the development of a gene index to the human genome: an assessment of the nature of high-throughput EST sequence data.
Genome Res. 1996 Sep;6(9):829-45
PMID: 8889550
-
Genome maps 7. The human transcript map. Wall chart.
Science. 1996 Oct 25;274(5287):547-62
PMID: 8928009
-
Sequence mapping by electronic PCR
Genome Res. 1997 May;7(5):541-50
PMID: 9149949
-
High throughput fingerprint analysis of large-insert clones.
Genome Res. 1997 Nov;7(11):1072-84
PMID: 9371743
-
Late-night thoughts on the sequence annotation problem.
Genome Res. 1998 Mar;8(3):168-9
PMID: 9521919
-
Base-calling of automated sequencer traces using phred. I. Accuracy assessment.
Genome Res. 1998 Mar;8(3):175-85
PMID: 9521921
-
Base-calling of automated sequencer traces using phred. II. Error probabilities.
Genome Res. 1998 Mar;8(3):186-94
PMID: 9521922
-
Estimation of errors in "raw" DNA sequences: a validation study.
Genome Res. 1998 Mar;8(3):251-9
PMID: 9521928
-
Alternative gene form discovery and candidate gene selection from gene indexing projects.
Genome Res. 1998 Mar;8(3):276-90
PMID: 9521931
-
Masquerading repeats: paralogous pitfalls of the human genome.
Genome Res. 1998 Aug;8(8):758-62
PMID: 9724321
-
A physical map of 30,000 human genes.
Science. 1998 Oct 23;282(5389):744-6
PMID: 9784132
-
The protein information resource (PIR).
Nucleic Acids Res. 2000 Jan 1;28(1):41-4
PMID: 10592177