Skip to main navigation Skip to search Skip to main content

Towards quantitative analysis of data intensive computing: A case study of Hadoop

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

In modern data centers, Hadoop has been widely used in perform data-intensive computation. Administrators of large scale hadoop clusters leverage statistical data collected at runtime to measure the efficiency of the cluster utilization. In this paper, we propose three statistical metrics - data locality ratio, load balance coefficient and access balance coefficient to quantify performance losses in data intensive applications. We evaluated our metrics using a large scale web click stream application running on a productive hadoop cluster at Tencent Inc.

Original languageEnglish
Title of host publicationProceedings of the 8th ACM International Conference on Autonomic Computing, ICAC 2011 and Co-located Workshops
Pages193-194
Number of pages2
DOIs
StatePublished - 2011
Externally publishedYes
Event8th ACM International Conference on Autonomic Computing, ICAC 2011 and Co-located Workshops - Karlsruhe, Germany
Duration: 14 Jun 201118 Jun 2011

Publication series

NameProceedings of the 8th ACM International Conference on Autonomic Computing, ICAC 2011 and Co-located Workshops

Conference

Conference8th ACM International Conference on Autonomic Computing, ICAC 2011 and Co-located Workshops
Country/TerritoryGermany
CityKarlsruhe
Period14/06/1118/06/11

Keywords

  • Hadoop
  • MapReduce
  • management
  • performance metrics

Fingerprint

Dive into the research topics of 'Towards quantitative analysis of data intensive computing: A case study of Hadoop'. Together they form a unique fingerprint.

Cite this