Title
Spatial computations over terabyte-sized images on hadoop platforms
Abstract
Our objective is to lower the barrier of executing spatial image computations in a computer cluster/cloud environment instead of in a desktop/laptop computing environment. We research two related problems encountered during an execution of spatial computations over terabyte-sized images using Apache Hadoop running on distributed computing resources. The two problems address (a) detection of spatial computations and their parameter estimation from a library of image processing functions, and (b) partitioning of image data for spatial image computations on Hadoop cluster/cloud computing platforms in order to minimize network data transfer. The first problem is solved by designing an iterative estimation methodology. The second problem is formulated as an optimization over three partitioning schemas (physical, logical without overlap and logical with overlap), and evaluated over several system configuration parameters. Our experimental results for the two problems demonstrate 100% accuracy in detecting spatial computations in the Java Advanced Imaging and ImageJ libraries, a speed-up of 5.36 between the default Hadoop physical partitioning and developed logical image partitioning with overlap, and 3.14 times faster execution of logical partitioning with overlap than the one without overlap. The novelty of our work is in designing an extension to Apache Hadoop to run a class of spatial image processing operations efficiently on a distributed computing resource.
Year
DOI
Venue
2014
10.1109/BigData.2014.7004311
BigData Conference
Keywords
Field
DocType
Java,data handling,image processing,optimisation,parallel processing,parameter estimation,Java advanced imaging,apache Hadoop,distributed computing resources,image processing functions,imageJ libraries,logical image partitioning,parameter estimation,Distributed computing,Hadoop,Image partition,Spatial image operations
Data mining,Data-intensive computing,Terabyte,Computer science,Parallel computing,Image processing,Estimation theory,Java,Computer cluster,Cloud computing,Computation
Conference
ISSN
Citations 
PageRank 
2639-1589
1
0.36
References 
Authors
5
4
Name
Order
Citations
PageRank
Peter Bajcsy113825.50
Phuong Nguyen210.36
Antoine Vandecreme371.79
Mary Brady43910.10