Home LiteratureArticle Details
PMID: 21932440 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Consequences of the discontinuation of the International Protein Index (IPI) database and its substitution by the UniProtKB "complete proteome" sets.

Proteomics ·Vol. 11 ·No. 22 ·2011-11-00 ·Pages 4434-8

Griss J, Martín M, O'Donovan C, Apweiler R, Hermjakob H, Vizcaíno JA

Abstract

The International Protein Index (IPI) database has been one of the most widely used protein databases in MS proteomics approaches. Recently, the closure of IPI in September 2011 was announced. Its recommended replacement is the new UniProt Knowledgebase (UniProtKB) "complete proteome" sets, launched in May 2011. Here, we analyze the consequences of IPI's discontinuation for human and mouse data, and the effect of its substitution with UniProtKB on two levels: (i) data already produced and (ii) newly performed experiments. To estimate the effect on existing data, we investigated how well IPI identifiers map to UniProtKB accessions. We found that 21% of human and 10% of mouse identifiers do not map to UniProtKB and would thus be "lost." To investigate the impact on new experiments, we compared the theoretical search space (i.e. the tryptic peptides) of both resources and found that it is decreased by 14.0% for human and 8.9% for mouse data through IPI's closure. An analysis on the experimental evidence for these "lost" peptides showed that the vast majority has not been identified in experiments available in the major proteomics repositories. It thus seems likely that the search space provided by UniProtKB is of higher quality than the one currently provided by IPI.

MeSH Terms
Animals Computational Biology/methods,organization & administration Database Management Systems Databases, Protein Humans Mice
Authors & Affiliations
6 authors, click to expand affiliations / ORCID
Griss Johannes
EMBL-European Bioinformatics Institute, Wellcome Trust Genome Campus, Hinxton, Cambridge, UK.
Martín María
O'Donovan Claire
Apweiler Rolf
Hermjakob Henning
Vizcaíno Juan Antonio
References (16)
16 references, click to expand
  1. Design, implementation and maintenance of a model organism database for Arabidopsis thaliana.
    Comp Funct Genomics. 2004;5(4):362-9 PMID: 18629167
  2. Published and perished? The influence of the searched protein database on the long-term storage of proteomics data.
    Mol Cell Proteomics. 2011 Sep;10(9):M111.008490 PMID: 21700957
  3. The Protein Identifier Cross-Referencing (PICR) service: reconciling protein identifiers across multiple source databases.
    BMC Bioinformatics. 2007 Oct 18;8:401 PMID: 17945017
  4. Large-scale database searching using tandem mass spectra: looking up the answer in the back of the book.
    Nat Methods. 2004 Dec;1(3):195-202 PMID: 15789030
  5. The human proteome project: current state and future direction.
    Mol Cell Proteomics. 2011 Jul;10(7):M111.009993 PMID: 21742803
  6. Ongoing and future developments at the Universal Protein Resource.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D214-9 PMID: 21051339
  7. A HUPO test sample study reveals common problems in mass spectrometry-based proteomics.
    Nat Methods. 2009 Jun;6(6):423-30 PMID: 19448641
  8. The International Protein Index: an integrated database for proteomics experiments.
    Proteomics. 2004 Jul;4(7):1985-8 PMID: 15221759
  9. PeptideAtlas: a resource for target selection for emerging targeted proteomics workflows.
    EMBO Rep. 2008 May;9(5):429-34 PMID: 18451766
  10. NCBI Reference Sequences: current status, policy and new initiatives.
    Nucleic Acids Res. 2009 Jan;37(Database issue):D32-6 PMID: 18927115
  11. The Proteomics Identifications database: 2010 update.
    Nucleic Acids Res. 2010 Jan;38(Database issue):D736-42 PMID: 19906717
  12. The consensus coding sequence (CCDS) project: Identifying a common protein-coding gene set for the human and mouse genomes.
    Genome Res. 2009 Jul;19(7):1316-23 PMID: 19498102
  13. Ensembl 2011.
    Nucleic Acids Res. 2011 Jan;39(Database issue):D800-6 PMID: 21045057
  14. The H-Invitational Database (H-InvDB), a comprehensive annotation resource for human genes and transcripts.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D793-9 PMID: 18089548
  15. The vertebrate genome annotation (Vega) database.
    Nucleic Acids Res. 2008 Jan;36(Database issue):D753-60 PMID: 18003653
  16. Open source system for analyzing, validating, and storing protein identification data.
    J Proteome Res. 2004 Nov-Dec;3(6):1234-42 PMID: 15595733
Article Info
Journal
Proteomics
Abbr.
Proteomics
ISSN
1615-9861
Published
2011-11-00
Epub
2011-00-17
Pages
4434-8
Language
English
Region
Germany
NLM ID
101092707
PMCID
PMC3556690
Subset
IM
Grants
Wellcome Trust · WT085949MA · United Kingdom
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]