DocumentCode
3761518
Title
Replication Management Framework for HDFS Based on Prediction Technique
Author
Dinh-Mao Bui;Thien Huynh-The;Sungyoung Lee;Bin Li;Jin Wang
Author_Institution
Comput. Eng. Dept., Kyung Hee Univ., Suwon, South Korea
fYear
2015
Firstpage
58
Lastpage
63
Abstract
The number of application based on Apache Hadoop is increasing dramatically due to the robustness and dynamic features of this system. At the heart of Apache Hadoop, the Hadoop File System (HDFS) provides the reliability, scalability and high availability to computation by applying a static replication strategy. However, because of the characteristics of parallel operations on the application layer, the accessing frequency for each data file in HDFS is totally different. Consequently, maintaining the same replicating mechanism for every data file might lead to bad effects on the performance. By rigorously considering the drawbacks of HDFS architecture, this paper proposes an approach to dynamically replicate the data file based on the predictive analysis. With the help of probability theory, the utilization of each data file can be predicted to create an individual replication strategy. Eventually, the data file can subsequently be replicated depending on its own access potential. Hence, this approach simultaneously improves the data locality while keeping the analogous redundancy of data storage in comparison with the default replicating scheme.
Keywords
"Big data","Monitoring","Training data","Kernel","Gaussian processes","Data mining","Pattern matching"
Publisher
ieee
Conference_Titel
Advanced Cloud and Big Data, 2015 Third International Conference on
Print_ISBN
978-1-4673-8537-4
Type
conf
DOI
10.1109/CBD.2015.19
Filename
7435453
Link To Document