Home LiteratureArticle Details
PMID: 27418748 Published · ppublish English Journal Article

DecoyPyrat: Fast Non-redundant Hybrid Decoy Sequence Generation for Large Scale Proteomics.

Journal of proteomics & bioinformatics ·Vol. 9 ·No. 6 ·2016-06-27 ·Pages 176-180

Wright JC, Choudhary JS

Abstract

Accurate statistical evaluation of sequence database peptide identifications from tandem mass spectra is essential in mass spectrometry based proteomics experiments. These statistics are dependent on accurately modelling random identifications. The target-decoy approach has risen to become the de facto approach to calculating FDR in proteomic datasets. The main principle of this approach is to search a set of decoy protein sequences that emulate the size and composition of the target protein sequences searched whilst not matching real proteins in the sample. To do this, it is commonplace to reverse or shuffle the proteins and peptides in the target database. However, these approaches have their drawbacks and limitations. A key confounding issue is the peptide redundancy between target and decoy databases leading to inaccurate FDR estimation. This inaccuracy is further amplified at the protein level and when searching large sequence databases such as those used for proteogenomics. Here, we present a unifying hybrid method to quickly and efficiently generate decoy sequences with minimal overlap between target and decoy peptides. We show that applying a reversed decoy approach can produce up to 5% peptide redundancy and many more additional peptides will have the exact same precursor mass as a target peptide. Our hybrid method addresses both these issues by first switching proteolytic cleavage sites with preceding amino acid, reversing the database and then shuffling any redundant sequences. This flexible hybrid method reduces the peptide overlap between target and decoy peptides to about 1% of peptides, making a more robust decoy model suitable for large search spaces. We also demonstrate the anti-conservative effect of redundant peptides on the calculation of q-values in mouse brain tissue data.

Keywords
Database searching FDR Python Sequence database Shotgun proteomics Target-decoy
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Wright James C
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, UK.
Choudhary Jyoti S
Wellcome Trust Sanger Institute, Wellcome Genome Campus, Hinxton, Cambridge CB10 1SA, UK.
References (18)
18 references, click to expand
  1. MS-GF+ makes progress towards a universal database search tool for proteomics.
    Nat Commun. 2014 Oct 31;5:5277 PMID: 25358478
  2. Empirical statistical model to estimate the accuracy of peptide identifications made by MS/MS and database search.
    Anal Chem. 2002 Oct 15;74(20):5383-92 PMID: 12403597
  3. Improvements to the percolator algorithm for Peptide identification from shotgun proteomics data sets.
    J Proteome Res. 2009 Jul;8(7):3737-45 PMID: 19385687
  4. Comparison of novel decoy database designs for optimizing protein identification searches using ABRF sPRG2006 standard MS/MS data sets.
    J Proteome Res. 2009 Apr;8(4):1782-91 PMID: 19714810
  5. Assigning significance to peptides identified by tandem mass spectrometry using decoy databases.
    J Proteome Res. 2008 Jan;7(1):29-34 PMID: 18067246
  6. Andromeda: a peptide search engine integrated into the MaxQuant environment.
    J Proteome Res. 2011 Apr 1;10(4):1794-805 PMID: 21254760
  7. Initial quantitative proteomic map of 28 mouse tissues using the SILAC mouse.
    Mol Cell Proteomics. 2013 Jun;12(6):1709-22 PMID: 23436904
  8. Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry.
    Nat Methods. 2007 Mar;4(3):207-14 PMID: 17327847
  9. A refined method to calculate false discovery rates for peptide identification using decoy databases.
    J Proteome Res. 2009 Apr;8(4):1792-6 PMID: 19714873
  10. Probability-based protein identification by searching sequence databases using mass spectrometry data.
    Electrophoresis. 1999 Dec;20(18):3551-67 PMID: 10612281
  11. Solution to Statistical Challenges in Proteomics Is More Statistics, Not Less.
    J Proteome Res. 2015 Oct 2;14(10):4099-103 PMID: 26257019
  12. Posterior error probabilities and false discovery rates: two sides of the same coin.
    J Proteome Res. 2008 Jan;7(1):40-4 PMID: 18052118
  13. Enhanced peptide identification by electron transfer dissociation using an improved Mascot Percolator.
    Mol Cell Proteomics. 2012 Aug;11(8):478-91 PMID: 22493177
  14. Proteogenomics: concepts, applications and computational strategies.
    Nat Methods. 2014 Nov;11(11):1114-25 PMID: 25357241
  15. An approach to correlate tandem mass spectral data of peptides with amino acid sequences in a protein database.
    J Am Soc Mass Spectrom. 1994 Nov;5(11):976-89 PMID: 24226387
  16. Creating reference gene annotation for the mouse C57BL6/J genome assembly.
    Mamm Genome. 2015 Oct;26(9-10):366-78 PMID: 26187010
  17. Semi-supervised learning for peptide identification from shotgun proteomics datasets.
    Nat Methods. 2007 Nov;4(11):923-5 PMID: 17952086
  18. Improving GENCODE reference gene annotation using a high-stringency proteogenomics workflow.
    Nat Commun. 2016 Jun 02;7:11778 PMID: 27250503
Article Info
Journal
Journal of proteomics & bioinformatics
Abbr.
J Proteomics Bioinform
ISSN
0974-276X
Published
2016-06-27
Pages
176-180
Language
English
Region
United States
NLM ID
101479045
PMCID
PMC4941923
Grants
Wellcome Trust · 098051 · United Kingdom
NHGRI NIH HHS · U41 HG007234 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]