Home LiteratureArticle Details
PMID: 19608599 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't

Protein identification false discovery rates for very large proteomics data sets generated by tandem mass spectrometry.

Molecular & cellular proteomics : MCP ·Vol. 8 ·No. 11 ·2009-11-00 ·Pages 2405-17

Reiter L, Claassen M, Schrimpf SP, Jovanovic M, Schmidt A, Buhmann JM, Hengartner MO, Aebersold R

Abstract

Comprehensive characterization of a proteome is a fundamental goal in proteomics. To achieve saturation coverage of a proteome or specific subproteome via tandem mass spectrometric identification of tryptic protein sample digests, proteomics data sets are growing dramatically in size and heterogeneity. The trend toward very large integrated data sets poses so far unsolved challenges to control the uncertainty of protein identifications going beyond well established confidence measures for peptide-spectrum matches. We present MAYU, a novel strategy that reliably estimates false discovery rates for protein identifications in large scale data sets. We validated and applied MAYU using various large proteomics data sets. The data show that the size of the data set has an important and previously underestimated impact on the reliability of protein identifications. We particularly found that protein false discovery rates are significantly elevated compared with those of peptide-spectrum matches. The function provided by MAYU is critical to control the quality of proteome data repositories and thereby to enhance any study relying on these data sources. The MAYU software is available as standalone software and also integrated into the Trans-Proteomic Pipeline.

MeSH Terms
Algorithms Animals Caenorhabditis elegans/metabolism Computational Biology/methods Databases, Protein False Positive Reactions Leptospira/metabolism Models, Statistical Peptides/chemistry Proteome Proteomics/methods Reproducibility of Results Schizosaccharomyces/metabolism Tandem Mass Spectrometry/methods
Chemicals
Peptides Proteome
Authors & Affiliations
8 authors, click to expand affiliations / ORCID
Reiter Lukas
Institute of Molecular Biology, University of Zurich, CH-8057 Zurich, Switzerland.
Claassen Manfred
Schrimpf Sabine P
Jovanovic Marko
Schmidt Alexander
Buhmann Joachim M
Hengartner Michael O
Aebersold Ruedi
References (41)
41 references, click to expand
  1. An integrated, directed mass spectrometric approach for in-depth characterization of complex peptide mixtures.
    Mol Cell Proteomics. 2008 Nov;7(11):2138-50 PMID: 18511481
  2. A high-quality catalog of the Drosophila melanogaster proteome.
    Nat Biotechnol. 2007 May;25(5):576-83 PMID: 17450130
  3. Genome-scale proteomics reveals Arabidopsis thaliana gene models and proteome dynamics.
    Science. 2008 May 16;320(5878):938-41 PMID: 18436743
  4. Empirical statistical model to estimate the accuracy of peptide identifications made by MS/MS and database search.
    Anal Chem. 2002 Oct 15;74(20):5383-92 PMID: 12403597
  5. An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database.
    J Am Soc Mass Spectrom. 1994 Nov;5(11):976-89 PMID: 24226387
  6. Probability-based validation of protein identifications using a modified SEQUEST algorithm.
    Anal Chem. 2002 Nov 1;74(21):5593-9 PMID: 12433093
  7. The Bioperl toolkit: Perl modules for the life sciences.
    Genome Res. 2002 Oct;12(10):1611-8 PMID: 12368254
  8. Development and validation of a spectral library searching method for peptide identification from MS/MS.
    Proteomics. 2007 Mar;7(5):655-67 PMID: 17295354
  9. Improving the success rate of proteome analysis by modeling protein-abundance distributions and experimental designs.
    Nat Biotechnol. 2007 Jun;25(6):651-5 PMID: 17557102
  10. Overview of the HUPO Plasma Proteome Project: results from the pilot phase with 35 collaborating laboratories and multiple analytical groups, generating a core dataset of 3020 proteins and a publicly-available database.
    Proteomics. 2005 Aug;5(13):3226-45 PMID: 16104056
  11. PRIDE: the proteomics identifications database.
    Proteomics. 2005 Aug;5(13):3537-45 PMID: 16041671
  12. Interpretation of shotgun proteomic data: the protein inference problem.
    Mol Cell Proteomics. 2005 Oct;4(10):1419-40 PMID: 16009968
  13. Open source system for analyzing, validating, and storing protein identification data.
    J Proteome Res. 2004 Nov-Dec;3(6):1234-42 PMID: 15595733
  14. Integration with the human genome of peptide sequences obtained by high-throughput mass spectrometry.
    Genome Biol. 2005;6(1):R9 PMID: 15642101
  15. Evaluation of multidimensional chromatography coupled with tandem mass spectrometry (LC/LC-MS/MS) for large-scale protein analysis: the yeast proteome.
    J Proteome Res. 2003 Jan-Feb;2(1):43-50 PMID: 12643542
  16. EBP, a program for protein identification using multiple tandem mass spectrometry datasets.
    Mol Cell Proteomics. 2007 Mar;6(3):527-36 PMID: 17164401
  17. Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry.
    Nat Methods. 2007 Mar;4(3):207-14 PMID: 17327847
  18. A uniform proteomics MS/MS analysis platform utilizing open XML file formats.
    Mol Syst Biol. 2005;1:2005.0017 PMID: 16729052
  19. Analysis of the Saccharomyces cerevisiae proteome with PeptideAtlas.
    Genome Biol. 2006;7(11):R106 PMID: 17101051
  20. Comprehensive mass-spectrometry-based proteome quantification of haploid versus diploid yeast.
    Nature. 2008 Oct 30;455(7217):1251-4 PMID: 18820680
  21. Coherent membrane supports for parallel microsynthesis and screening of bioactive peptides.
    Biopolymers. 2000;55(3):188-206 PMID: 11074414
  22. Mass spectrometry-based proteomics.
    Nature. 2003 Mar 13;422(6928):198-207 PMID: 12634793
  23. A mammalian organelle map by protein correlation profiling.
    Cell. 2006 Apr 7;125(1):187-99 PMID: 16615899
  24. Qscore: an algorithm for evaluating SEQUEST database search results.
    J Am Soc Mass Spectrom. 2002 Apr;13(4):378-86 PMID: 11951976
  25. Peptide arrays on cellulose support: SPOT synthesis, a time and cost efficient method for synthesis of large numbers of peptides in a parallel and addressable fashion.
    Nat Protoc. 2007;2(6):1333-49 PMID: 17545971
  26. Assigning significance to peptides identified by tandem mass spectrometry using decoy databases.
    J Proteome Res. 2008 Jan;7(1):29-34 PMID: 18067246
  27. Data management and preliminary data analysis in the pilot phase of the HUPO Plasma Proteome Project.
    Proteomics. 2005 Aug;5(13):3246-61 PMID: 16104057
  28. Scoring proteomes with proteotypic peptide probes.
    Nat Rev Mol Cell Biol. 2005 Jul;6(7):577-83 PMID: 15957003
  29. Large-scale analysis of the yeast proteome by multidimensional protein identification technology.
    Nat Biotechnol. 2001 Mar;19(3):242-7 PMID: 11231557
  30. Comparative functional analysis of the Caenorhabditis elegans and Drosophila melanogaster proteomes.
    PLoS Biol. 2009 Mar 3;7(3):e48 PMID: 19260763
  31. Computational prediction of proteotypic peptides for quantitative proteomics.
    Nat Biotechnol. 2007 Jan;25(1):125-31 PMID: 17195840
  32. A Heuristic method for assigning a false-discovery rate for protein identifications from Mascot database search results.
    Mol Cell Proteomics. 2005 Jun;4(6):762-72 PMID: 15703444
  33. Analysis and validation of proteomic data generated by tandem mass spectrometry.
    Nat Methods. 2007 Oct;4(10):787-97 PMID: 17901868
  34. Deterministic protein inference for shotgun proteomics data provides new insights into Arabidopsis pollen development and function.
    Genome Res. 2009 Oct;19(10):1786-800 PMID: 19546170
  35. Using annotated peptide mass spectrum libraries for protein identification.
    J Proteome Res. 2006 Aug;5(8):1843-9 PMID: 16889405
  36. Chemical substructure identification by mass spectral library searching.
    J Am Soc Mass Spectrom. 1995 Aug;6(8):644-55 PMID: 24214391
  37. A model for random sampling and estimation of relative protein abundance in shotgun proteomics.
    Anal Chem. 2004 Jul 15;76(14):4193-201 PMID: 15253663
  38. A method for the comprehensive proteomic analysis of membrane proteins.
    Nat Biotechnol. 2003 May;21(5):532-8 PMID: 12692561
  39. Sperm chromatin proteomics identifies evolutionarily conserved fertility factors.
    Nature. 2006 Sep 7;443(7107):101-5 PMID: 16943775
  40. A statistical model for identifying proteins by tandem mass spectrometry.
    Anal Chem. 2003 Sep 1;75(17):4646-58 PMID: 14632076
  41. What does it mean to identify a protein in proteomics?
    Trends Biochem Sci. 2002 Feb;27(2):74-8 PMID: 11852244
Article Info
Journal
Molecular & cellular proteomics : MCP
Abbr.
Mol Cell Proteomics
ISSN
1535-9484
Published
2009-11-00
Epub
2009-00-16
Pages
2405-17
Language
English
Region
United States
NLM ID
101125647
PMCID
PMC2773710
Subset
IM
Grants
NHLBI NIH HHS · N01HV28179 · United States
NHLBI NIH HHS · N01-HV-28179 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]