Big data text-oriented benchmark creation for Hadoop - Citegraph

Paper Info

Title
Big data text-oriented benchmark creation for Hadoop

Abstract
Massive-scale Big Data analytics is representative of a new class of workloads that justifies a rethinking of how computing systems should be optimized. This paper addresses the need for a set of benchmarks that system designers can use to measure the quality of their designs and that customers can use to evaluate competing systems offerings with respect to commonly performed text-oriented workflows in Hadoop™. Additions are needed to existing benchmarks such as HiBench in terms of both scale and relevance. We describe a methodology for creating a petascale data-size text-oriented benchmark that includes representative Big Data workflows and can be used to test total system performance, with demands balanced across storage, network, and computation. Creating such a benchmark requires meeting unique challenges associated with the data size and its often unstructured nature. To be useful, the benchmark also needs to be sufficiently generic to be accepted by the community at large. Here, we focus on a text-oriented Hadoop workflow that consists of three common tasks: categorizing text documents, identifying significant documents within each category, and analyzing significant documents for new topic creation.

Year	DOI	Venue
2013	10.1147/JRD.2013.2240732	IBM Journal of Research and Development
Keywords	Field	DocType
text-oriented workflows,benchmark creation,petascale data-size text-oriented benchmark,significant document,new class,computing system,representative big data workflows,new topic creation,big data,system designer,massive-scale big data analytics,text-oriented hadoop workflow	Data science,Computer science,Petascale computing,Big data,Workflow,Computing systems,Computation	Journal
Volume	Issue	ISSN
57	3-4	0018-8646
Citations	PageRank	References
15	1.26	4
Authors
5

Authors (5 rows)

Cited by (15 rows)

References (4 rows)

Name	Order	Citations	PageRank
Gattiker, A.	1	15	1.26
F. H. Gebara	2	18	1.89
H. P. Hofstee	3	507	54.92
J. D. Hayes	4	15	1.26
A. Hylick	5	15	1.26

1