Title
A Snapshot of the OWL Web
Abstract
Tool development for and empirical experimentation in OWL ontology engineering require a wide variety of suitable ontologies as input for testing and evaluation purposes and detailed characterisations of real ontologies. Empirical activities often resort to (somewhat arbitrarily) hand curated corpora available on the web, such as the NCBO BioPortal and the TONES Repository, or manually selected sets of well-known ontologies. Findings of surveys and results of benchmarking activities may be biased, even heavily, towards these datasets. Sampling from a large corpus of ontologies, on the other hand, may lead to more representative results. Current large scale repositories and web crawls are mostly uncurated and suffer from duplication, small and (for many purposes) uninteresting ontology files, and contain large numbers of ontology versions, variants, and facets, and therefore do not lend themselves to random sampling. In this paper, we survey ontologies as they exist on the web and describe the creation of a corpus of OWL DL ontologies using strategies such as web crawling, various forms of de-duplications and manual cleaning, which allows random sampling of ontologies for a variety of empirical applications.
Year
DOI
Venue
2013
10.1007/978-3-642-41335-3_21
International Semantic Web Conference (1)
Keywords
Field
DocType
corpus,empirical methods,ontology engineering,owl
Ontology (information science),Ontology engineering,Data mining,World Wide Web,Process ontology,Computer science,Open Biomedical Ontologies,IDEF5,OWL-S,Ontology components,Database,Web Ontology Language
Conference
Volume
ISSN
Citations 
8218
0302-9743
20
PageRank 
References 
Authors
0.94
14
3
Name
Order
Citations
PageRank
Nicolas Matentzoglu17914.97
Samantha Bail21068.03
Bijan Parsia36146429.53