Abstract | ||
---|---|---|
Rank fusion is a powerful technique that allows multiple sources of information to be combined into a single result set. Query variations covering the same information need represent one way in which different sources of information might arise. However, when implemented in the obvious manner, fusion over query variations is not cost-effective, at odds with the usual web-search requirement for strict per-query efficiency guarantees. In this work, we propose a novel solution to query fusion by splitting the computation into two parts: one phase that is carried out offline, to generate pre-computed centroid answers for queries addressing broadly similar information needs, and then a second online phase that uses the corresponding topic centroid to compute a result page for each query. To achieve this, we make use of score-based fusion algorithms whose costs can be amortized via the pre-processing step and that can then be efficiently combined during subsequent per-query re-ranking operations. Experimental results using the ClueWeb12B collection and the UQV100 query variations demonstrate that centroid-based approaches allow improved retrieval effectiveness at little or no loss in query throughput or latency and within reasonable pre-processing requirements. We additionally show that queries that do not match any of the pre-computed clusters can be accurately identified and efficiently processed in our proposed ranking pipeline.
|
Year | DOI | Venue |
---|---|---|
2019 | 10.1145/3345001 | ACM Transactions on Information Systems |
Keywords | Field | DocType |
Rank fusion,dynamic pruning,effectiveness,efficiency,experimentation | Information retrieval,Computer science,Boosting (machine learning) | Journal |
Volume | Issue | ISSN |
37 | 4 | 1046-8188 |
Citations | PageRank | References |
2 | 0.38 | 0 |
Authors | ||
4 |
Name | Order | Citations | PageRank |
---|---|---|---|
Rodger Benham | 1 | 7 | 3.19 |
Joel Mackenzie | 2 | 44 | 10.36 |
Alistair Moffat | 3 | 5913 | 728.91 |
Shane Culpepper | 4 | 519 | 47.52 |