Home LiteratureArticle Details
PMID: 15608268 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

HOPPSIGEN: a database of human and mouse processed pseudogenes.

Nucleic acids research ·Vol. 33 ·No. Database issue ·2005-01-01 ·Pages D59-66

Khelifi A, Adel K, Duret L, Laurent D, Mouchiroud D, Dominique M

Abstract

Processed pseudogenes result from reverse transcribed mRNAs. In general, because processed pseudogenes lack promoters, they are no longer functional from the moment they are inserted into the genome. Subsequently, they freely accumulate substitutions, insertions and deletions. Moreover, the ancestral structure of processed pseudogenes could be easily inferred using the sequence of their functional homologous genes. Owing to these characteristics, processed pseudogenes represent good neutral markers for studying genome evolution. Recently, there is an increasing interest for these markers, particularly to help gene prediction in the field of genome annotation, functional genomics and genome evolution analysis (patterns of substitution). For these reasons, we have developed a method to annotate processed pseudogenes in complete genomes. To make them useful to different fields of research, we stored them in a nucleic acid database after having annotated them. In this work, we screened both mouse and human complete genomes from ENSEMBL to find processed pseudogenes generated from functional genes with introns. We used a conservative method to detect processed pseudogenes in order to minimize the rate of false positive sequences. Within processed pseudogenes, some are still having a conserved open reading frame and some have overlapping gene locations. We designated as retroelements all reverse transcribed sequences and more strictly, we designated as processed pseudogenes, all retroelements not falling in the two former categories (having a conserved open reading or overlapping gene locations). We annotated 5823 retroelements (5206 processed pseudogenes) in the human genome and 3934 (3428 processed pseudogenes) in the mouse genome. Compared to previous estimations, the total number of processed pseudogenes was underestimated but the aim of this procedure was to generate a high-quality dataset. To facilitate the use of processed pseudogenes in studying genome structure and evolution, DNA sequences from processed pseudogenes, and their functional reverse transcribed homologs, are now stored in a nucleic acid database, HOPPSIGEN. HOPPSIGEN can be browsed on the PBIL (Pole Bioinformatique Lyonnais) World Wide Web server (http://pbil.univ-lyon1.fr/) or fully downloaded for local installation.

MeSH Terms
Animals Databases, Nucleic Acid Humans Internet Mice Pseudogenes Retroelements Reverse Transcription
Chemicals
Retroelements
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Khelifi Adel
Laboratoire de Biométrie et Biologie Evolutive, UMR CNRS 5558, Université Claude Bernard-Lyon 1, 43 bd. du 11 Novembre 1918, 69622 Villeurbanne Cedex, France. [email protected]
Adel Khelifi
Duret Laurent
Laurent Duret
Mouchiroud Dominique
Dominique Mouchiroud
References (30)
30 references, click to expand
  1. Vertebrate pseudogenes.
    FEBS Lett. 2000 Feb 25;468(2-3):109-14 PMID: 10692568
  2. Isochores result from mutation not selection.
    Nature. 1999 Jul 1;400(6739):30-1 PMID: 10403245
  3. Human LINE retrotransposons generate processed pseudogenes.
    Nat Genet. 2000 Apr;24(4):363-7 PMID: 10742098
  4. Nature and structure of human genes that generate retropseudogenes.
    Genome Res. 2000 May;10(5):672-8 PMID: 10810090
  5. Analysis of expressed sequence tags indicates 35,000 human genes.
    Nat Genet. 2000 Jun;25(2):232-4 PMID: 10835644
  6. Characterization and repeat analysis of the compact genome of the freshwater pufferfish Tetraodon nigroviridis.
    Genome Res. 2000 Jul;10(7):939-49 PMID: 10899143
  7. Initial sequencing and analysis of the human genome.
    Nature. 2001 Feb 15;409(6822):860-921 PMID: 11237011
  8. Molecular fossils in the human genome: identification and analysis of the pseudogenes in chromosomes 21 and 22.
    Genome Res. 2002 Feb;12(2):272-80 PMID: 11827946
  9. Identification and analysis of over 2000 ribosomal protein pseudogenes in the human genome.
    Genome Res. 2002 Oct;12(10):1466-82 PMID: 12368239
  10. Initial sequencing and comparative analysis of the mouse genome.
    Nature. 2002 Dec 5;420(6915):520-62 PMID: 12466850
  11. Length distribution of long interspersed nucleotide elements (LINEs) and processed pseudogenes of human endogenous retroviruses: implications for retrotransposition and pseudogene detection.
    Gene. 2002 Oct 30;300(1-2):189-94 PMID: 12468100
  12. The SWISS-PROT protein knowledgebase and its supplement TrEMBL in 2003.
    Nucleic Acids Res. 2003 Jan 1;31(1):365-70 PMID: 12520024
  13. Human-mouse alignments with BLASTZ.
    Genome Res. 2003 Jan;13(1):103-7 PMID: 12529312
  14. Molecular biology: Complicity of gene and pseudogene.
    Nature. 2003 May 1;423(6935):26-8 PMID: 12721611
  15. An expressed pseudogene regulates the messenger-RNA stability of its homologous coding gene.
    Nature. 2003 May 1;423(6935):91-6 PMID: 12721631
  16. Integrated databanks access and sequence/structure analysis services at the PBIL.
    Nucleic Acids Res. 2003 Jul 1;31(13):3393-9 PMID: 12824334
  17. Whole-genome screening indicates a possible burst of formation of processed pseudogenes and Alu repeats by particular L1 subfamilies in ancestral primates.
    Genome Biol. 2003;4(11):R74 PMID: 14611660
  18. Millions of years of evolution preserved: a comprehensive catalog of the processed pseudogenes in the human genome.
    Genome Res. 2003 Dec;13(12):2541-58 PMID: 14656962
  19. A genome-wide survey of human pseudogenes.
    Genome Res. 2003 Dec;13(12):2559-67 PMID: 14656963
  20. The EMBL Nucleotide Sequence Database.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D27-30 PMID: 14681351
  21. Ensembl 2004.
    Nucleic Acids Res. 2004 Jan 1;32(Database issue):D468-70 PMID: 14681459
  22. Comparative analysis of processed pseudogenes in the mouse and human genomes.
    Trends Genet. 2004 Feb;20(2):62-7 PMID: 14746985
  23. A simple method for estimating evolutionary rates of base substitutions through comparative studies of nucleotide sequences.
    J Mol Evol. 1980 Dec;16(2):111-20 PMID: 7463489
  24. Processed pseudogenes: characteristics and evolution.
    Annu Rev Genet. 1985;19:253-72 PMID: 3909943
  25. The neighbor-joining method: a new method for reconstructing phylogenetic trees.
    Mol Biol Evol. 1987 Jul;4(4):406-25 PMID: 3447015
  26. ACNUC--a portable retrieval system for nucleic acid sequence databases: logical and physical designs and usage.
    Comput Appl Biosci. 1985 Sep;1(3):167-72 PMID: 3880341
  27. HOVERGEN: a database of homologous vertebrate genes.
    Nucleic Acids Res. 1994 Jun 25;22(12):2360-5 PMID: 8036164
  28. CLUSTAL W: improving the sensitivity of progressive multiple sequence alignment through sequence weighting, position-specific gap penalties and weight matrix choice.
    Nucleic Acids Res. 1994 Nov 11;22(22):4673-80 PMID: 7984417
  29. On-line tools for sequence retrieval and multivariate statistics in molecular biology.
    Comput Appl Biosci. 1996 Feb;12(1):63-9 PMID: 8670621
  30. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs.
    Nucleic Acids Res. 1997 Sep 1;25(17):3389-402 PMID: 9254694
Article Info
Journal
Nucleic acids research
Abbr.
Nucleic Acids Res
ISSN
1362-4962
Published
2005-01-01
Pages
D59-66
Language
English
Region
England
NLM ID
0411011
PMCID
PMC540038
Subset
IM
Corrections
ErratumIn
-
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]