Abstract | ||
---|---|---|
This work addresses the problem of detecting novel sentences from an incoming stream of text data, by studying the performance of different novelty metrics, and proposing a mixed metric that is able to adapt to different performance requirements. Existing novelty metrics can be divided into two types, symmetric and asymmetric, based on whether the ordering of sentences is taken into account. After a comparative study of several different novelty metrics, we observe complementary behavior in the two types of metrics. This finding motivates a new framework of novelty measurement, i.e. the mixture of both symmetric and asymmetric metrics. This new framework of novelty measurement performs superiorly under different performance requirements varying from high-precision to high-recall as well as for data with different percentages of novel sentences. Because it does not require any prior information, the new metric is very suitable for real-time knowledge base applications such as novelty mining systems where no training data is available beforehand. |
Year | DOI | Venue |
---|---|---|
2010 | 10.1016/j.ins.2010.02.020 | Inf. Sci. |
Keywords | Field | DocType |
sentence-level novelty mining,novel sentence,different percentage,novelty mining system,new metric,novelty metrics,novelty measurement,new framework,different performance requirement,different novelty metrics,asymmetric metrics,knowledge base,information retrieval,real time | Training set,Novelty detection,Computer science,Artificial intelligence,Novelty,Knowledge base,Sentence,Machine learning | Journal |
Volume | Issue | ISSN |
180 | 12 | 0020-0255 |
Citations | PageRank | References |
18 | 0.67 | 17 |
Authors | ||
3 |
Name | Order | Citations | PageRank |
---|---|---|---|
Flora S. Tsai | 1 | 352 | 23.96 |
Wenyin Tang | 2 | 90 | 7.19 |
Kap Luk Chan | 3 | 1039 | 77.99 |