In-depth analysis of protein inference algorithms using multiple search engines and well-defined metrics

Audain, Enrique; Uszkoreit, Julian; Sachsenberg, Timo; Pfeuffer, Julianus; Liang, Xiao; Hermjakob, Henning; Sanchez, Aniel; Eisenacher, Martin; Reinert, Knut; Tabb, David L.; Kohlbacher, Oliver; Perez-Riverol, Yasset

In-depth analysis of protein inference algorithms using multiple search engines and well-defined metrics

Mark

Audain, Enrique ; Uszkoreit, Julian ; Sachsenberg, Timo ; Pfeuffer, Julianus ; Liang, Xiao ; Hermjakob, Henning ; Sanchez, Aniel ^LU ; Eisenacher, Martin ; Reinert, Knut and Tabb, David L. , et al. (2017) In Journal of Proteomics 150. p.170-182

Abstract: In mass spectrometry-based shotgun proteomics, protein identifications are usually the desired result. However, most of the analytical methods are based on the identification of reliable peptides and not the direct identification of intact proteins. Thus, assembling peptides identified from tandem mass spectra into a list of proteins, referred to as protein inference, is a critical step in proteomics research. Currently, different protein inference algorithms and tools are available for the proteomics community. Here, we evaluated five software tools for protein inference (PIA, ProteinProphet, Fido, ProteinLP, MSBayesPro) using three popular database search engines: Mascot, X!Tandem, and MS-GF +. All the algorithms were evaluated using... (More); In mass spectrometry-based shotgun proteomics, protein identifications are usually the desired result. However, most of the analytical methods are based on the identification of reliable peptides and not the direct identification of intact proteins. Thus, assembling peptides identified from tandem mass spectra into a list of proteins, referred to as protein inference, is a critical step in proteomics research. Currently, different protein inference algorithms and tools are available for the proteomics community. Here, we evaluated five software tools for protein inference (PIA, ProteinProphet, Fido, ProteinLP, MSBayesPro) using three popular database search engines: Mascot, X!Tandem, and MS-GF +. All the algorithms were evaluated using a highly customizable KNIME workflow using four different public datasets with varying complexities (different sample preparation, species and analytical instruments). We defined a set of quality control metrics to evaluate the performance of each combination of search engines, protein inference algorithm, and parameters on each dataset. We show that the results for complex samples vary not only regarding the actual numbers of reported protein groups but also concerning the actual composition of groups. Furthermore, the robustness of reported proteins when using databases of differing complexities is strongly dependant on the applied inference algorithm. Finally, merging the identifications of multiple search engines does not necessarily increase the number of reported proteins, but does increase the number of peptides per protein and thus can generally be recommended. Significance Protein inference is one of the major challenges in MS-based proteomics nowadays. Currently, there are a vast number of protein inference algorithms and implementations available for the proteomics community. Protein assembly impacts in the final results of the research, the quantitation values and the final claims in the research manuscript. Even though protein inference is a crucial step in proteomics data analysis, a comprehensive evaluation of the many different inference methods has never been performed. Previously Journal of proteomics has published multiple studies about other benchmark of bioinformatics algorithms (PMID: 26585461; PMID: 22728601) in proteomics studies making clear the importance of those studies for the proteomics community and the journal audience. This manuscript presents a new bioinformatics solution based on the KNIME/OpenMS platform that aims at providing a fair comparison of protein inference algorithms (https://github.com/KNIME-OMICS). Six different algorithms - ProteinProphet, MSBayesPro, ProteinLP, Fido and PIA- were evaluated using the highly customizable workflow on four public datasets with varying complexities. Five popular database search engines Mascot, X!Tandem, MS-GF + and combinations thereof were evaluated for every protein inference tool. In total > 186 proteins lists were analyzed and carefully compare using three metrics for quality assessments of the protein inference results: 1) the numbers of reported proteins, 2) peptides per protein, and the 3) number of uniquely reported proteins per inference method, to address the quality of each inference method. We also examined how many proteins were reported by choosing each combination of search engines, protein inference algorithms and parameters on each dataset. The results show that using 1) PIA or Fido seems to be a good choice when studying the results of the analyzed workflow, regarding not only the reported proteins and the high-quality identifications, but also the required runtime. 2) Merging the identifications of multiple search engines gives almost always more confident results and increases the number of peptides per protein group. 3) The usage of databases containing not only the canonical, but also known isoforms of proteins has a small impact on the number of reported proteins. The detection of specific isoforms could, concerning the question behind the study, compensate for slightly shorter reports using the parsimonious reports. 4) The current workflow can be easily extended to support new algorithms and search engine combinations.
(Less)

Please use this url to cite or link to this publication: https://lup.lub.lu.se/record/630b899a-f06b-4354-85ce-6a0d1441b8ca

author

Audain, Enrique ; Uszkoreit, Julian ; Sachsenberg, Timo ; Pfeuffer, Julianus ; Liang, Xiao ; Hermjakob, Henning ; Sanchez, Aniel ^LU ; Eisenacher, Martin ; Reinert, Knut and Tabb, David L. , et al. (More)

Audain, Enrique ; Uszkoreit, Julian ; Sachsenberg, Timo ; Pfeuffer, Julianus ; Liang, Xiao ; Hermjakob, Henning ; Sanchez, Aniel ^LU ; Eisenacher, Martin ; Reinert, Knut ; Tabb, David L. ; Kohlbacher, Oliver and Perez-Riverol, Yasset (Less)

organization

Department of Translational Medicine

publishing date

2017-01-06

type

Contribution to journal

publication status

published

subject

Bioinformatics (Computational Biology)

keywords

Algorithms, Bioinformatics, Mass spectrometry, Protein inference

in

Journal of Proteomics

volume

150

pages

13 pages

publisher

Elsevier

external identifiers

pmid:27498275
wos:000390621400016
scopus:84988843652

ISSN

1874-3919

DOI

10.1016/j.jprot.2016.08.002

language

English

LU publication?

yes

id

630b899a-f06b-4354-85ce-6a0d1441b8ca

date added to LUP

2016-10-18 14:33:02

date last changed

2024-04-05 08:22:44

@article{630b899a-f06b-4354-85ce-6a0d1441b8ca,
  abstract     = {{<p>In mass spectrometry-based shotgun proteomics, protein identifications are usually the desired result. However, most of the analytical methods are based on the identification of reliable peptides and not the direct identification of intact proteins. Thus, assembling peptides identified from tandem mass spectra into a list of proteins, referred to as protein inference, is a critical step in proteomics research. Currently, different protein inference algorithms and tools are available for the proteomics community. Here, we evaluated five software tools for protein inference (PIA, ProteinProphet, Fido, ProteinLP, MSBayesPro) using three popular database search engines: Mascot, X!Tandem, and MS-GF +. All the algorithms were evaluated using a highly customizable KNIME workflow using four different public datasets with varying complexities (different sample preparation, species and analytical instruments). We defined a set of quality control metrics to evaluate the performance of each combination of search engines, protein inference algorithm, and parameters on each dataset. We show that the results for complex samples vary not only regarding the actual numbers of reported protein groups but also concerning the actual composition of groups. Furthermore, the robustness of reported proteins when using databases of differing complexities is strongly dependant on the applied inference algorithm. Finally, merging the identifications of multiple search engines does not necessarily increase the number of reported proteins, but does increase the number of peptides per protein and thus can generally be recommended. Significance Protein inference is one of the major challenges in MS-based proteomics nowadays. Currently, there are a vast number of protein inference algorithms and implementations available for the proteomics community. Protein assembly impacts in the final results of the research, the quantitation values and the final claims in the research manuscript. Even though protein inference is a crucial step in proteomics data analysis, a comprehensive evaluation of the many different inference methods has never been performed. Previously Journal of proteomics has published multiple studies about other benchmark of bioinformatics algorithms (PMID: 26585461; PMID: 22728601) in proteomics studies making clear the importance of those studies for the proteomics community and the journal audience. This manuscript presents a new bioinformatics solution based on the KNIME/OpenMS platform that aims at providing a fair comparison of protein inference algorithms (https://github.com/KNIME-OMICS). Six different algorithms - ProteinProphet, MSBayesPro, ProteinLP, Fido and PIA- were evaluated using the highly customizable workflow on four public datasets with varying complexities. Five popular database search engines Mascot, X!Tandem, MS-GF + and combinations thereof were evaluated for every protein inference tool. In total &gt; 186 proteins lists were analyzed and carefully compare using three metrics for quality assessments of the protein inference results: 1) the numbers of reported proteins, 2) peptides per protein, and the 3) number of uniquely reported proteins per inference method, to address the quality of each inference method. We also examined how many proteins were reported by choosing each combination of search engines, protein inference algorithms and parameters on each dataset. The results show that using 1) PIA or Fido seems to be a good choice when studying the results of the analyzed workflow, regarding not only the reported proteins and the high-quality identifications, but also the required runtime. 2) Merging the identifications of multiple search engines gives almost always more confident results and increases the number of peptides per protein group. 3) The usage of databases containing not only the canonical, but also known isoforms of proteins has a small impact on the number of reported proteins. The detection of specific isoforms could, concerning the question behind the study, compensate for slightly shorter reports using the parsimonious reports. 4) The current workflow can be easily extended to support new algorithms and search engine combinations.</p>}},
  author       = {{Audain, Enrique and Uszkoreit, Julian and Sachsenberg, Timo and Pfeuffer, Julianus and Liang, Xiao and Hermjakob, Henning and Sanchez, Aniel and Eisenacher, Martin and Reinert, Knut and Tabb, David L. and Kohlbacher, Oliver and Perez-Riverol, Yasset}},
  issn         = {{1874-3919}},
  keywords     = {{Algorithms; Bioinformatics; Mass spectrometry; Protein inference}},
  language     = {{eng}},
  month        = {{01}},
  pages        = {{170--182}},
  publisher    = {{Elsevier}},
  series       = {{Journal of Proteomics}},
  title        = {{In-depth analysis of protein inference algorithms using multiple search engines and well-defined metrics}},
  url          = {{https://lup.lub.lu.se/search/files/20360849/15724392.pdf}},
  doi          = {{10.1016/j.jprot.2016.08.002}},
  volume       = {{150}},
  year         = {{2017}},
}

Lund University Publications

LUND UNIVERSITY LIBRARIES

In-depth analysis of protein inference algorithms using multiple search engines and well-defined metrics