Resumen
Large-scale cloud datacenters often experience reduced performance and service outage. Due to the inherent complexity, heterogeneity, and multitenant architecture of these datacenters, applications (i.e., jobs and tasks) running on them are susceptible to various types of failures. In this paper, we first characterize the application failures in Google cluster trace and then propose a prediction model which can forecast the termination status of a task. Then, we introduce a task scheduler that dynamically reschedules tasks based on the predicted results. This proactive fault-tolerant scheduler improves system reliability and ensures timely execution of the applications. Simulation results show that our scheduler reduces makespan and failure rates of tasks substantially while balancing load at the same time. Moreover, early prediction along with quick scheduling adjustment improves overall resource utilization and reduces resource wastage.
| Idioma original | English |
|---|---|
| Título de la publicación alojada | Proceedings - 6th IEEE International Conference on Cyber Security and Cloud Computing, CSCloud 2019 and 5th IEEE International Conference on Edge Computing and Scalable Cloud, EdgeCom 2019 |
| Editores | Meikang Qiu |
| Páginas | 1-6 |
| Número de páginas | 6 |
| ISBN (versión digital) | 9781728116600 |
| DOI | |
| Estado | Published - jun 2019 |
| Evento | 6th IEEE International Conference on Cyber Security and Cloud Computing and 5th IEEE International Conference on Edge Computing and Scalable Cloud, CSCloud/EdgeCom 2019 - Paris, France Duración: jun 21 2019 → jun 23 2019 |
Serie de la publicación
| Nombre | Proceedings - 6th IEEE International Conference on Cyber Security and Cloud Computing, CSCloud 2019 and 5th IEEE International Conference on Edge Computing and Scalable Cloud, EdgeCom 2019 |
|---|
Conference
| Conference | 6th IEEE International Conference on Cyber Security and Cloud Computing and 5th IEEE International Conference on Edge Computing and Scalable Cloud, CSCloud/EdgeCom 2019 |
|---|---|
| País/Territorio | France |
| Ciudad | Paris |
| Período | 6/21/19 → 6/23/19 |
Nota bibliográfica
Publisher Copyright:© 2019 IEEE.
ASJC Scopus subject areas
- Computer Networks and Communications
- Hardware and Architecture
- Safety, Risk, Reliability and Quality
Huella
Profundice en los temas de investigación de 'FaCS: Toward a Fault-Tolerant Cloud Scheduler Leveraging Long Short-Term Memory Network'. En conjunto forman una huella única.Citar esto
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver