Title
HPVdb: a Data Mining System for Knowledge Discovery in Human Papillomavirus with Applications in T cell Immunology and Vaccinology
Abstract
High-risk human papillomaviruses (HPVs) are the causes of many cancers, including cervical, anal, vulvar, vaginal, penile and oropharyngeal. To facilitate diagnosis, prognosis and characterization of these cancers, it is necessary to make full use of the immunological data on HPV available through publications, technical reports and databases. These data vary in granularity, quality and complexity. The extraction of knowledge from the vast amount of immunological data using data mining techniques remains a challenging task. To support integration of data and knowledge in virology and vaccinology, we developed a framework called KB-builder to streamline the development and deployment of web-accessible immunological knowledge systems. The framework consists of seven major functional modules, each facilitating a specific aspect of the knowledgebase construction process. Using KB-builder, we constructed the Human Papillomavirus T cell Antigen Database (HPVdb). It contains 2781 curated antigen entries of antigenic proteins derived from 18 genotypes of high-risk HPV and 18 genotypes of low-risk HPV. The HPVdb also catalogs 191 verified T cell epitopes and 45 verified human leukocyte antigen (HLA) ligands. Primary amino acid sequences of HPV antigens were collected and annotated from the UniProtKB. T cell epitopes and HLA ligands were collected from data mining of scientific literature and databases. The data were subject to extensive quality control (redundancy elimination, error detection and vocabulary consolidation). A set of computational tools for an in-depth analysis, such as sequence comparison using BLAST search, multiple alignments of antigens, classification of HPV types based on cancer risk, T cell epitope/HLA ligand visualization, T cell epitope/HLA ligand conservation analysis and sequence variability analysis, has been integrated within the HPVdb. Predicted Class I and Class II HLA binding peptides for 15 common HLA alleles are included in this database as putative targets. HPVdb is a knowledge-based system that integrates curated data and information with tailored analysis tools to facilitate data mining for HPV vaccinology and immunology. To our best knowledge, HPVdb is a unique data source providing a comprehensive list of HPV antigens and peptides. Database URL: http://cvc.dfci.harvard.edu/hpv/.
Year
DOI
Venue
2014
10.1145/2506583.2512360
Database : the journal of biological databases and curation
Keywords
Field
DocType
knowledge discovery,hla ligands,class ii hla,data mining system,high-risk hpv,cell epitopes,low-risk hpv,hpv type,cell immunology,common hla allele,hla ligand visualization,human papillomavirus,curated data,hpv antigens
Epitope,Data mining,Papillomaviridae,Antigen,Biology,Immunology,UniProt,Vaccination,Bioinformatics,Human leukocyte antigen,Papillomavirus Vaccines,T cell
Journal
Volume
ISSN
Citations 
2014
1758-0463
1
PageRank 
References 
Authors
0.35
4
6
Name
Order
Citations
PageRank
G.L. Zhang120816.91
Angelika B Riemer210.35
Derin B. Keskin3111.07
Lou Chitkushev4319.07
Ellis L Reinherz5956.10
Vladimir Brusic655163.37