An Analysis of Traces from a Production MapReduce Cluster 论文

2010引用 319
Cloud Computing and Resource ManagementSoftware System Performance and ReliabilityIoT and Edge/Fog Computing

详细信息

发表日期
2010-01-01
发表年份
2010

关键词

Cloud Computing and Resource ManagementSoftware System Performance and ReliabilityIoT and Edge/Fog Computing

摘要

MapReduce is a programming paradigm for parallel processing that is increasingly being used for data-intensive applications in cloud computing environments. An understanding of the characteristics of workloads running in MapReduce environments benefits both the service providers in the cloud and users: the service provider can use this knowledge to make better scheduling decisions, while the user can learn what aspects of their jobs impact performance. This paper analyzes 10-months of MapReduce logs from the M45 supercomputing cluster which Yahoo! made freely available to select universities for academic research. We characterize resource utilization patterns, job patterns, and sources of failures. We use an instance-based learning technique that exploits temporal locality to predict job completion times from historical data and identify potential performance problems in our dataset.

相关技术

暂无数据

相关事件

暂无数据

相关文章

暂无数据