Dependency-Aware Rollback and Checkpoint-Restart for Distributed   Task-Based Runtimes

Kiril Dichev; Herbert Jordan; Konstantinos Tovletoglou and; Thomas Heller; Dimitrios S. Nikolopoulos; Georgios Karakonstantis and; Charles Gillan

arXiv:1705.10208·cs.DC·May 30, 2017·2 cites

Dependency-Aware Rollback and Checkpoint-Restart for Distributed Task-Based Runtimes

Kiril Dichev, Herbert Jordan, Konstantinos Tovletoglou and, Thomas Heller, Dimitrios S. Nikolopoulos, Georgios Karakonstantis and, Charles Gillan

PDF

Open Access

TL;DR

This paper introduces a dependency-aware rollback technique for task-based runtimes that reduces task cancellation and recomputation during failures, leading to faster recovery and execution times.

Contribution

It proposes a novel dependency-aware rollback method combined with a recursive decomposition-based checkpoint/restart technique for task-based runtimes.

Findings

01

Dependency-aware rollback reduces task cancellation during failures.

02

The combined approach improves recovery speed in simulations.

03

Faster overall execution time observed despite no guarantee of reduced total execution time.

Abstract

With the increase in compute nodes in large compute platforms, a proportional increase in node failures will follow. Many application-based checkpoint/restart (C/R) techniques have been proposed for MPI applications to target the reduced mean time between failures. However, rollback as part of the recovery remains a dominant cost even in highly optimised MPI applications employing C/R techniques. Continuing execution past a checkpoint (that is, reducing rollback) is possible in message-passing runtimes, but extremely complex to design and implement. Our work focuses on task-based runtimes, where task dependencies are explicit and message passing is implicit. We see an opportunity for reducing rollback for such runtimes: we explore task dependencies in the rollback, which we call dependency-aware rollback. We also design a new C/R technique, which is influenced by recursive decomposition…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDistributed systems and fault tolerance · Distributed and Parallel Computing Systems · Parallel Computing and Optimization Techniques