Title
RHEEMix in the data jungle: a cost-based optimizer for cross-platform systems
Abstract
Data analytics are moving beyond the limits of a single platform. In this paper, we present the cost-based optimizer of Rheem, an open-source cross-platform system that copes with these new requirements. The optimizer allocates the subtasks of data analytic tasks to the most suitable platforms. Our main contributions are: (i) a mechanism based on graph transformations to explore alternative execution strategies; (ii) a novel graph-based approach to determine efficient data movement plans among subtasks and platforms; and (iii) an efficient plan enumeration algorithm, based on a novel enumeration algebra. We extensively evaluate our optimizer under diverse real tasks. We show that our optimizer can perform tasks more than one order of magnitude faster when using multiple platforms than when using a single platform.
Year
DOI
Venue
2020
10.1007/s00778-020-00612-x
The VLDB Journal
Keywords
DocType
Volume
Cross-platform, Polystore, Query optimization, Data processing
Journal
29
Issue
ISSN
Citations 
6
1066-8888
1
PageRank 
References 
Authors
0.43
10
6
Name
Order
Citations
PageRank
Sebastian Kruse1518.03
Zoi Kaoudi221518.55
Bertty Contreras-Rojas312.46
Sanjay Chawla41372105.09
Felix Naumann51900174.92
Jorge-arnulfo Quiané-ruiz698661.02