Title
A Workflow Application for Parallel Processing of Big Data from an Internet Portal.
Abstract
The paper presents a workflow application for efficient parallel processing of data downloaded from an Internet portal. The workflow partitions input files into subdirectories which are further split for parallel processing by services installed on distinct computer nodes. This way, analysis of the first ready sub-directories can start fast and is handled by services implemented as parallel multithreaded applications using multiple cores of modern CPUs. The goal is to assess achievable speed-ups and determine which factors influence scalability and to what degree. Data processing services were implemented for assessment of context (positive or negative) in which the given keyword appears in a document. The testbed application used these services to determine how a particular brand was recognized by either authors of articles or readers in comments in a specific Internet portal focused on new technologies. Obtained execution times as well as speed-ups are presented for data sets of various sizes along with discussion on how factors such as load imbalance and memory/disk bottlenecks limit performance.
Year
DOI
Venue
2014
10.1016/j.procs.2014.05.045
Procedia Computer Science
Keywords
Field
DocType
parallel data processing,parallel performance,Internet,parallel workflow application
Data mining,Workflow technology,Computer science,Workflow application,Workflow engine,Workflow management system,Workflow,Big data,Operating system,Database,The Internet,Scalability
Conference
Volume
ISSN
Citations 
29
1877-0509
1
PageRank 
References 
Authors
0.34
10
1
Name
Order
Citations
PageRank
Pawel Czarnul112121.11