A Symbolic Regression Based Scoring System Improving Peptide Identifications for MS Amanda - Citegraph

Paper Info

Title
A Symbolic Regression Based Scoring System Improving Peptide Identifications for MS Amanda

Abstract
Peptide search engines are algorithms that are able to identify peptides (i.e., short proteins or parts of proteins) from mass spectra of biological samples. These identification algorithms report the best matching peptide for a given spectrum and a score that represents the quality of the match; usually, the higher this score, the higher is the reliability of the respective match. In order to estimate the specificity and sensitivity of search engines, sets of target sequences are given to the identification algorithm as well as so-called decoy sequences that are randomly created or scrambled versions of real sequences; decoy sequences should be assigned low scores whereas target sequences should be assigned high scores. In this paper we present an approach based on symbolic regression (using genetic programming) that helps to distinguish between target and decoy matches. On the basis of features calculated for matched sequences and using the information on the original sequence set (target or decoy) we learn mathematical models that calculate updated scores. As an alternative to this white box modeling approach we also use a black box modeling method, namely random forests. As we show in the empirical section of this paper, this approach leads to scores that increase the number of reliably identified samples that are originally scored using the MS Amanda identification algorithm for high resolution as well as for low resolution mass spectra.

Year	DOI	Venue
2015	10.1145/2739482.2768509	GECCO (Companion)
Field	DocType	Citations
Black box (phreaking),Search engine,Decoy,Computer science,White box,Genetic programming,Artificial intelligence,Random forest,Mathematical model,Symbolic regression,Machine learning	Conference	0
PageRank	References	Authors
0.34	2	5

Authors (5 rows)

Cited by (0 rows)

References (2 rows)

Name	Order	Citations	PageRank
Viktoria Dorfer	1	28	2.49
Sergey Maltsev	2	0	0.34
Stephan Dreiseitl	3	338	34.80
Karl Mechtler	4	27	1.99
Stephan M. Winkler	5	140	22.90

1