SVM-Fold: a tool for discriminative multi-class protein fold and superfamily recognition. - Citegraph

Paper Info

Title
SVM-Fold: a tool for discriminative multi-class protein fold and superfamily recognition.

Abstract
Predicting a protein's structural class from its amino acid sequence is a fundamental problem in computational biology. Much recent work has focused on developing new representations for protein sequences, called string kernels, for use with support vector machine (SVM) classifiers. However, while some of these approaches exhibit state-of-the-art performance at the binary protein classification problem, i.e. discriminating between a particular protein class and all other classes, few of these studies have addressed the real problem of multi-class superfamily or fold recognition. Moreover, there are only limited software tools and systems for SVM-based protein classification available to the bioinformatics community.We present a new multi-class SVM-based protein fold and superfamily recognition system and web server called SVM-Fold, which can be found at http://svm-fold.c2b2.columbia.edu. Our system uses an efficient implementation of a state-of-the-art string kernel for sequence profiles, called the profile kernel, where the underlying feature representation is a histogram of inexact matching k-mer frequencies. We also employ a novel machine learning approach to solve the difficult multi-class problem of classifying a sequence of amino acids into one of many known protein structural classes. Binary one-vs-the-rest SVM classifiers that are trained to recognize individual structural classes yield prediction scores that are not comparable, so that standard "one-vs-all" classification fails to perform well. Moreover, SVMs for classes at different levels of the protein structural hierarchy may make useful predictions, but one-vs-all does not try to combine these multiple predictions. To deal with these problems, our method learns relative weights between one-vs-the-rest classifiers and encodes information about the protein structural hierarchy for multi-class prediction. In large-scale benchmark results based on the SCOP database, our code weighting approach significantly improves on the standard one-vs-all method for both the superfamily and fold prediction in the remote homology setting and on the fold recognition problem. Moreover, our code weight learning algorithm strongly outperforms nearest-neighbor methods based on PSI-BLAST in terms of prediction accuracy on every structure classification problem we consider.By combining state-of-the-art SVM kernel methods with a novel multi-class algorithm, the SVM-Fold system delivers efficient and accurate protein fold and superfamily recognition.

Year	DOI	Venue
2007	10.1186/1471-2105-8-S4-S2	BMC Bioinformatics
Keywords	Field	DocType
machine learning,string kernel,protein structure,algorithms,fold recognition,amino acid,support vector machine,amino acid sequence,protein folding,computational biology,protein sequence,bioinformatics,microarrays,nearest neighbor method	Sequence alignment,Protein folding,Pattern recognition,Computer science,Support vector machine,Threading (protein sequence),Artificial intelligence,Linear discriminant analysis,Bioinformatics,String kernel,Discriminative model,Protein structure	Journal
Volume	Issue	ISSN
8 Suppl 4	S-4	1471-2105
Citations	PageRank	References
47	1.42	15
Authors
6

Authors (6 rows)

Cited by (47 rows)

References (15 rows)

Name	Order	Citations	PageRank
Iain Melvin	1	129	5.04
Eugene Ie	2	237	14.45
Rui Kuang	3	484	31.16
Jason Weston	4	13068	805.30
William Stafford Noble	5	2907	203.56
Christina Leslie	6	1389	77.99

1