Logo image
ALBERT: An automatic learning based execution and resource management system for optimizing Hadoop workload in clouds
期刊文章   同儕審查

ALBERT: An automatic learning based execution and resource management system for optimizing Hadoop workload in clouds

Chen-Chun Chen, Kai-Siang Wang, Yu-Tung HsiaoJerry Chou
Journal of Parallel and Distributed Computing, 卷.168, 頁碼.45-56
10/2022

摘要

Data analytic Deep learning Job scheduling Optimization Time prediction Software Theoretical Computer Science Hardware and Architecture Computer Networks and Communications Artificial Intelligence
Hadoop is a popular computing framework designed to deliver timely and cost-effective data processing on a large cluster of commodity machines. It relieves the burden of the programmers dealing with distributed programming, and an ecosystem of Big Data solutions has developed around it. However, Hadoop's job execution time can greatly depend on its runtime configurations and resource selections. Given the more than 100 job configuration settings provided by Hadoop, and diverse resource instance options in a cloud or virtualized computing environment, running Hadoop jobs still requires a substantial amount of expertise and experience. To address this challenge, we apply a deep neural network to predict Hadoop's job time based on historical execution data, and propose optimization methods to reduce job execution time and cost. The results show that our prediction method achieves almost 90% time prediction accuracy and clearly outperforms three other state-of-the-art regression-based prediction methods. Based on the time prediction, our proposed configuration search method and job scheduling algorithm successfully shorten the execution time of a single Hadoop job by more than a factor of 2 and reduce the time of processing a batch of Hadoop jobs by 40%∼65%.

相關連結

指標

1 檢視次數

詳細資料

Logo image