Home LiteratureArticle Details
PMID: 11337484 Published · ppublish English Comparative Study Journal Article Research Support, Non-U.S. Gov't

Assembly, annotation, and integration of UNIGENE clusters into the human genome draft.

Genome research ·Vol. 11 ·No. 5 ·2001-05-00 ·Pages 904-18

Zhuo D, Zhao WD, Wright FA, Yang HY, Wang JP, Sears R, Baer T, Kwon DH, Gordon D, Gibbs S, Dai D, Yang Q, Spitzner J, Krahe R, Stredney D, Stutz A, Yuan B

Abstract

The recent release of the first draft of the human genome provides an unprecedented opportunity to integrate human genes and their functions in a complete positional context. However, at least three significant technical hurdles remain: first, to assemble a complete and nonredundant human transcript index; second, to accurately place the individual transcript indices on the human genome; and third, to functionally annotate all human genes. Here, we report the extension of the UNIGENE database through the assembly of its sequence clusters into nonredundant sequence contigs. Each resulting consensus was aligned to the human genome draft. A unique location for each transcript within the human genome was determined by the integration of the restriction fingerprint, assembled genomic contig, and radiation hybrid (RH) maps. A total of 59,500 UNIGENE clusters were mapped on the basis of at least three independent criteria as compared with the 30,000 human genes/ESTs currently mapped in Genemap'99. Finally, the extension of the human transcript consensus in this study enabled a greater number of putative functional assignments than the 11,000 annotated entries in UNIGENE. This study reports a draft physical map with annotations for a majority of the human transcripts, called the Human Index of Nonredundant Transcripts (HINT). Such information can be immediately applied to the discovery of new genes and the identification of candidate genes for positional cloning.

MeSH Terms
Alleles Alternative Splicing/genetics Computational Biology/methods Consensus Sequence/genetics Databases, Factual Genes/genetics Genome, Human Human Genome Project Humans Multigene Family/genetics Sequence Alignment/methods
Authors & Affiliations
17 authors, click to expand affiliations / ORCID
Zhuo D
Bioinformatics Group, James Cancer Hospital and Solove Research Institute, The Ohio State University, Columbus, Ohio 43210, USA.
Zhao W D
Wright F A
Yang H Y
Wang J P
Sears R
Baer T
Kwon D H
Gordon D
Gibbs S
Dai D
Yang Q
Spitzner J
Krahe R
Stredney D
Stutz A
Yuan B
References (31)
31 references, click to expand
  1. A comprehensive approach to clustering of expressed human gene sequence: the sequence tag alignment and consensus knowledge base.
    Genome Res. 1999 Nov;9(11):1143-55 PMID: 10568754
  2. Reliable identification of large numbers of candidate SNPs from public EST data.
    Nat Genet. 1999 Mar;21(3):323-5 PMID: 10080189
  3. The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000.
    Nucleic Acids Res. 2000 Jan 1;28(1):45-8 PMID: 10592178
  4. The TIGR gene indices: reconstruction and representation of expressed gene sequences.
    Nucleic Acids Res. 2000 Jan 1;28(1):141-5 PMID: 10592205
  5. The Pfam protein families database.
    Nucleic Acids Res. 2000 Jan 1;28(1):263-6 PMID: 10592242
  6. Frequent alternative splicing of human genes.
    Genome Res. 1999 Dec;9(12):1288-93 PMID: 10613851
  7. Representation of functional information in the SWISS-PROT data bank.
    Bioinformatics. 1999 Dec;15(12):1066-7 PMID: 10746001
  8. Analysis of expressed sequence tags indicates 35,000 human genes.
    Nat Genet. 2000 Jun;25(2):232-4 PMID: 10835644
  9. Gene index analysis of the human genome estimates approximately 120,000 genes.
    Nat Genet. 2000 Jun;25(2):239-40 PMID: 10835646
  10. Repeat polymorphisms within gene regions: phenotypic and evolutionary implications.
    Am J Hum Genet. 2000 Aug;67(2):345-56 PMID: 10889045
  11. Genome-wide analysis of single-nucleotide polymorphisms in human expressed sequences.
    Nat Genet. 2000 Oct;26(2):233-6 PMID: 11017085
  12. Basic local alignment search tool.
    J Mol Biol. 1990 Oct 5;215(3):403-10 PMID: 2231712
  13. Two acetyl-CoA acetyltransferase genes located in the t-complex region of mouse chromosome 17 partially overlap the Tcp-1 and Tcp-1x genes.
    Genomics. 1993 Nov;18(2):195-8 PMID: 7904580
  14. Gene discovery in dbEST.
    Science. 1994 Sep 30;265(5181):1993-4 PMID: 8091218
  15. False association of human ESTs.
    Nat Genet. 1994 Dec;8(4):321-2 PMID: 7894479
  16. ESTablishing a human transcript map.
    Nat Genet. 1995 Aug;10(4):369-71 PMID: 7670480
  17. Initial assessment of human gene diversity and expression patterns based upon 83 million nucleotides of cDNA sequence.
    Nature. 1995 Sep 28;377(6547 Suppl):3-174 PMID: 7566098
  18. The Genexpress Index: a resource for gene discovery and the genic map of the human genome.
    Genome Res. 1995 Oct;5(3):272-304 PMID: 8593614
  19. Generation and analysis of 280,000 human expressed sequence tags.
    Genome Res. 1996 Sep;6(9):807-28 PMID: 8889549
  20. Toward the development of a gene index to the human genome: an assessment of the nature of high-throughput EST sequence data.
    Genome Res. 1996 Sep;6(9):829-45 PMID: 8889550
  21. Genome maps 7. The human transcript map. Wall chart.
    Science. 1996 Oct 25;274(5287):547-62 PMID: 8928009
  22. Sequence mapping by electronic PCR
    Genome Res. 1997 May;7(5):541-50 PMID: 9149949
  23. High throughput fingerprint analysis of large-insert clones.
    Genome Res. 1997 Nov;7(11):1072-84 PMID: 9371743
  24. Late-night thoughts on the sequence annotation problem.
    Genome Res. 1998 Mar;8(3):168-9 PMID: 9521919
  25. Base-calling of automated sequencer traces using phred. I. Accuracy assessment.
    Genome Res. 1998 Mar;8(3):175-85 PMID: 9521921
  26. Base-calling of automated sequencer traces using phred. II. Error probabilities.
    Genome Res. 1998 Mar;8(3):186-94 PMID: 9521922
  27. Estimation of errors in "raw" DNA sequences: a validation study.
    Genome Res. 1998 Mar;8(3):251-9 PMID: 9521928
  28. Alternative gene form discovery and candidate gene selection from gene indexing projects.
    Genome Res. 1998 Mar;8(3):276-90 PMID: 9521931
  29. Masquerading repeats: paralogous pitfalls of the human genome.
    Genome Res. 1998 Aug;8(8):758-62 PMID: 9724321
  30. A physical map of 30,000 human genes.
    Science. 1998 Oct 23;282(5389):744-6 PMID: 9784132
  31. The protein information resource (PIR).
    Nucleic Acids Res. 2000 Jan 1;28(1):41-4 PMID: 10592177
Article Info
Journal
Genome research
Abbr.
Genome Res
ISSN
1088-9051
Published
2001-05-00
Pages
904-18
Language
English
Region
United States
NLM ID
9518021
PMCID
PMC311045
Subset
IM
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]