Phrase Query Optimization on Inverted Indexes - Citegraph

Paper Info

Title
Phrase Query Optimization on Inverted Indexes

Abstract
Phrase queries are a key functionality of modern search engines. Beyond that, they increasingly serve as an important building block for applications such as entity-oriented search, text analytics, and plagiarism detection. Processing phrase queries is costly, though, since positional information has to be kept in the index and all words, including stopwords, need to be considered. We consider an augmented inverted index that indexes selected variable-length multi-word sequences in addition to single words. We study how arbitrary phrase queries can be processed efficiently on such an augmented inverted index. We show that the underlying optimization problem is NP-hard in the general case and describe an exact exponential algorithm and an approximation algorithm to its solution. Experiments on ClueWeb09 and The New York Times with different real-world query workloads examine the practical performance of our methods.

Year	DOI	Venue
2014	10.1145/2661829.2661928	CIKM
Keywords	Field	DocType
multi-word indexing,phrase queries,query optimization,search process	Inverted index,Query optimization,Approximation algorithm,Search engine,Phrase search,Information retrieval,Plagiarism detection,Computer science,Phrase,Optimization problem	Conference
Citations	PageRank	References
1	0.35	15
Authors
4

Authors (4 rows)

Cited by (1 rows)

References (15 rows)

Name	Order	Citations	PageRank
Avishek Anand	1	14	5.10
Ida Mele	2	1	0.35
Srikanta Bedathur	3	607	43.23
Klaus Berberich	4	1271	68.96

1