Home LiteratureArticle Details
PMID: 21876204 Published · ppublish English Journal Article Research Support, N.I.H., Extramural Research Support, Non-U.S. Gov't Research Support, U.S. Gov't, Non-P.H.S.

iProphet: multi-level integrative analysis of shotgun proteomic data improves peptide and protein identification rates and error estimates.

Molecular & cellular proteomics : MCP ·Vol. 10 ·No. 12 ·2011-12-00 ·Pages M111.007690

Shteynberg D, Deutsch EW, Lam H, Eng JK, Sun Z, Tasman N, Mendoza L, Moritz RL, Aebersold R, Nesvizhskii AI

Abstract

The combination of tandem mass spectrometry and sequence database searching is the method of choice for the identification of peptides and the mapping of proteomes. Over the last several years, the volume of data generated in proteomic studies has increased dramatically, which challenges the computational approaches previously developed for these data. Furthermore, a multitude of search engines have been developed that identify different, overlapping subsets of the sample peptides from a particular set of tandem mass spectrometry spectra. We present iProphet, the new addition to the widely used open-source suite of proteomic data analysis tools Trans-Proteomics Pipeline. Applied in tandem with PeptideProphet, it provides more accurate representation of the multilevel nature of shotgun proteomic data. iProphet combines the evidence from multiple identifications of the same peptide sequences across different spectra, experiments, precursor ion charge states, and modified states. It also allows accurate and effective integration of the results from multiple database search engines applied to the same data. The use of iProphet in the Trans-Proteomics Pipeline increases the number of correctly identified peptides at a constant false discovery rate as compared with both PeptideProphet and another state-of-the-art tool Percolator. As the main outcome, iProphet permits the calculation of accurate posterior probabilities and false discovery rate estimates at the level of sequence identical peptide identifications, which in turn leads to more accurate probability estimates at the protein level. Fully integrated with the Trans-Proteomics Pipeline, it supports all commonly used MS instruments, search engines, and computer platforms. The performance of iProphet is demonstrated on two publicly available data sets: data from a human whole cell lysate proteome profiling experiment representative of typical proteomic data sets, and from a set of Streptococcus pyogenes experiments more representative of organism-specific composite data sets.

MeSH Terms
Algorithms Amino Acid Sequence Data Interpretation, Statistical Humans Jurkat Cells Peptide Fragments/chemistry Probability Proteome/chemistry Proteomics Search Engine Software Streptococcus pyogenes Tandem Mass Spectrometry
Chemicals
Peptide Fragments Proteome
Authors & Affiliations
10 authors, click to expand affiliations / ORCID
Shteynberg David
Institute for Systems Biology, Seattle, WA, USA.
Deutsch Eric W
Lam Henry
Eng Jimmy K
Sun Zhi
Tasman Natalie
Mendoza Luis
Moritz Robert L
Aebersold Ruedi
Nesvizhskii Alexey I
References (57)
57 references, click to expand
  1. False discovery rates and related statistical concepts in mass spectrometry-based proteomics.
    J Proteome Res. 2008 Jan;7(1):47-50 PMID: 18067251
  2. Empirical statistical model to estimate the accuracy of peptide identifications made by MS/MS and database search.
    Anal Chem. 2002 Oct 15;74(20):5383-92 PMID: 12403597
  3. Statistical calibration of the SEQUEST XCorr function.
    J Proteome Res. 2009 Apr;8(4):2106-13 PMID: 19275164
  4. A method for reducing the time required to match protein sequences with tandem mass spectra.
    Rapid Commun Mass Spectrom. 2003;17(20):2310-6 PMID: 14558131
  5. ProbID: a probabilistic algorithm to identify peptides through sequence database searching using tandem mass spectral data.
    Proteomics. 2002 Oct;2(10):1406-12 PMID: 12422357
  6. Modes of inference for evaluating the confidence of peptide identifications.
    J Proteome Res. 2008 Jan;7(1):35-9 PMID: 18067248
  7. Decoy methods for assessing false positives and false discovery rates in shotgun proteomics.
    Anal Chem. 2009 Jan 1;81(1):146-59 PMID: 19061407
  8. Proteome-wide cellular protein concentrations of the human pathogen Leptospira interrogans.
    Nature. 2009 Aug 6;460(7256):762-5 PMID: 19606093
  9. Open mass spectrometry search algorithm.
    J Proteome Res. 2004 Sep-Oct;3(5):958-64 PMID: 15473683
  10. Improving sensitivity in proteome studies by analysis of false discovery rates for multiple search engines.
    Proteomics. 2009 Mar;9(5):1220-9 PMID: 19253293
  11. OLAV: towards high-throughput tandem mass spectrometry data identification.
    Proteomics. 2003 Aug;3(8):1454-63 PMID: 12923771
  12. Statistical validation of peptide identifications in large-scale proteomics using the target-decoy database search strategy and flexible mixture modeling.
    J Proteome Res. 2008 Jan;7(1):286-92 PMID: 18078310
  13. Interpretation of large-scale quantitative shotgun proteomic profiles for biomarker discovery.
    Curr Opin Mol Ther. 2008 Jun;10(3):231-42 PMID: 18535930
  14. Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry.
    Nat Methods. 2007 Mar;4(3):207-14 PMID: 17327847
  15. A uniform proteomics MS/MS analysis platform utilizing open XML file formats.
    Mol Syst Biol. 2005;1:2005.0017 PMID: 16729052
  16. Getting started in computational mass spectrometry-based proteomics.
    PLoS Comput Biol. 2009 May;5(5):e1000366 PMID: 19492072
  17. A method for assessing the statistical significance of mass spectrometry-based protein identifications using general scoring schemes.
    Anal Chem. 2003 Feb 15;75(4):768-74 PMID: 12622365
  18. A high-quality catalog of the Drosophila melanogaster proteome.
    Nat Biotechnol. 2007 May;25(5):576-83 PMID: 17450130
  19. Global survey of human T leukemic cells by integrating proteomics and transcriptomics profiling.
    Mol Cell Proteomics. 2007 Aug;6(8):1343-53 PMID: 17519225
  20. Probabilistic assembly of human protein interaction networks from label-free quantitative proteomics.
    Proc Natl Acad Sci U S A. 2008 Feb 5;105(5):1454-9 PMID: 18218781
  21. The need for guidelines in publication of peptide and protein identification data: Working Group on Publication Guidelines for Peptide and Protein Identification Data.
    Mol Cell Proteomics. 2004 Jun;3(6):531-3 PMID: 15075378
  22. Improving sensitivity by probabilistically combining results from multiple MS/MS search methodologies.
    J Proteome Res. 2008 Jan;7(1):245-53 PMID: 18173222
  23. Genome-scale proteomics reveals Arabidopsis thaliana gene models and proteome dynamics.
    Science. 2008 May 16;320(5878):938-41 PMID: 18436743
  24. Optimized peptide separation and identification for mass spectrometry based proteomics via free-flow electrophoresis.
    J Proteome Res. 2006 Sep;5(9):2241-9 PMID: 16944936
  25. Spectral probabilities and generating functions of tandem mass spectra: a strike against decoy databases.
    J Proteome Res. 2008 Aug;7(8):3354-63 PMID: 18597511
  26. Probability-based protein identification by searching sequence databases using mass spectrometry data.
    Electrophoresis. 1999 Dec;20(18):3551-67 PMID: 10612281
  27. An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database.
    J Am Soc Mass Spectrom. 1994 Nov;5(11):976-89 PMID: 24226387
  28. General framework for developing and evaluating database scoring algorithms using the TANDEM search engine.
    Bioinformatics. 2006 Nov 15;22(22):2830-2 PMID: 16877754
  29. Comprehensive mass-spectrometry-based proteome quantification of haploid versus diploid yeast.
    Nature. 2008 Oct 30;455(7217):1251-4 PMID: 18820680
  30. Mass spectrometry-based proteomics.
    Nature. 2003 Mar 13;422(6928):198-207 PMID: 12634793
  31. Comparison of novel decoy database designs for optimizing protein identification searches using ABRF sPRG2006 standard MS/MS data sets.
    J Proteome Res. 2009 Apr;8(4):1782-91 PMID: 19714810
  32. Recent developments in proteome informatics for mass spectrometry analysis.
    Comb Chem High Throughput Screen. 2009 Feb;12(2):194-202 PMID: 19199887
  33. A hypergeometric probability model for protein identification and validation using tandem mass spectral data and protein sequence databases.
    Anal Chem. 2003 Aug 1;75(15):3792-8 PMID: 14572045
  34. Defining the human deubiquitinating enzyme interaction landscape.
    Cell. 2009 Jul 23;138(2):389-403 PMID: 19615732
  35. Calibrating E-values for MS2 database search methods.
    Biol Direct. 2007 Nov 05;2:26 PMID: 17983478
  36. Semi-supervised learning for peptide identification from shotgun proteomics datasets.
    Nat Methods. 2007 Nov;4(11):923-5 PMID: 17952086
  37. Enhancing peptide identification confidence by combining search methods.
    J Proteome Res. 2008 Aug;7(8):3102-13 PMID: 18558733
  38. A common open representation of mass spectrometry data and its application to proteomics research.
    Nat Biotechnol. 2004 Nov;22(11):1459-66 PMID: 15529173
  39. Analysis and validation of proteomic data generated by tandem mass spectrometry.
    Nat Methods. 2007 Oct;4(10):787-97 PMID: 17901868
  40. A refined method to calculate false discovery rates for peptide identification using decoy databases.
    J Proteome Res. 2009 Apr;8(4):1792-6 PMID: 19714873
  41. Development and validation of a spectral library searching method for peptide identification from MS/MS.
    Proteomics. 2007 Mar;7(5):655-67 PMID: 17295354
  42. Data analysis and bioinformatics tools for tandem mass spectrometry in proteomics.
    Physiol Genomics. 2008 Mar 14;33(1):18-25 PMID: 18212004
  43. A global protein kinase and phosphatase interaction network in yeast.
    Science. 2010 May 21;328(5981):1043-6 PMID: 20489023
  44. Adaptive discriminant function analysis and reranking of MS/MS database search results for improved peptide identification in shotgun proteomics.
    J Proteome Res. 2008 Nov;7(11):4878-89 PMID: 18788775
  45. A guided tour of the Trans-Proteomic Pipeline.
    Proteomics. 2010 Mar;10(6):1150-9 PMID: 20101611
  46. Posterior error probabilities and false discovery rates: two sides of the same coin.
    J Proteome Res. 2008 Jan;7(1):40-4 PMID: 18052118
  47. Interpretation of shotgun proteomic data: the protein inference problem.
    Mol Cell Proteomics. 2005 Oct;4(10):1419-40 PMID: 16009968
  48. Protein identification false discovery rates for very large proteomics data sets generated by tandem mass spectrometry.
    Mol Cell Proteomics. 2009 Nov;8(11):2405-17 PMID: 19608599
  49. MyriMatch: highly accurate tandem mass spectral peptide identification by multivariate hypergeometric analysis.
    J Proteome Res. 2007 Feb;6(2):654-61 PMID: 17269722
  50. Proteomics by mass spectrometry: approaches, advances, and applications.
    Annu Rev Biomed Eng. 2009;11:49-79 PMID: 19400705
  51. A statistical model for identifying proteins by tandem mass spectrometry.
    Anal Chem. 2003 Sep 1;75(17):4646-58 PMID: 14632076
  52. Integration with the human genome of peptide sequences obtained by high-throughput mass spectrometry.
    Genome Biol. 2005;6(1):R9 PMID: 15642101
  53. Semisupervised model-based validation of peptide identifications in mass spectrometry-based proteomics.
    J Proteome Res. 2008 Jan;7(1):254-65 PMID: 18159924
  54. A survey of computational methods and error rate estimation procedures for peptide and protein identification in shotgun proteomics.
    J Proteomics. 2010 Oct 10;73(11):2092-123 PMID: 20816881
  55. InsPecT: identification of posttranslationally modified peptides from tandem mass spectra.
    Anal Chem. 2005 Jul 15;77(14):4626-39 PMID: 16013882
  56. PeptideAtlas: a resource for target selection for emerging targeted proteomics workflows.
    EMBO Rep. 2008 May;9(5):429-34 PMID: 18451766
  57. Quantitative mass spectrometry in proteomics: a critical review.
    Anal Bioanal Chem. 2007 Oct;389(4):1017-31 PMID: 17668192
Article Info
Journal
Molecular & cellular proteomics : MCP
Abbr.
Mol Cell Proteomics
ISSN
1535-9484
Published
2011-12-00
Epub
2011-00-29
Pages
M111.007690
Language
English
Region
United States
NLM ID
101125647
PMCID
PMC3237071
Subset
IM
Grants
NHGRI NIH HHS · RC2 HG005805 · United States
NHLBI NIH HHS · N01-HV-28179 · United States
NIGMS NIH HHS · R01 GM087221 · United States
NCI NIH HHS · R01 CA126239 · United States
NHLBI NIH HHS · N01HV28179 · United States
NIGMS NIH HHS · R01 GM094231 · United States
NIGMS NIH HHS · P50 GM076547 · United States
NIGMS NIH HHS · PM50 GM076547 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]