Scalable Fault-Tolerant MapReduce

Demian Hespe; Lukas H\"ubner; Charel Mercatoris; Peter Sanders

arXiv:2411.16255·cs.DC·November 26, 2024

Scalable Fault-Tolerant MapReduce

Demian Hespe, Lukas H\"ubner, Charel Mercatoris, Peter Sanders

PDF

Open Access

TL;DR

This paper introduces a scalable, fault-tolerant MapReduce approach that minimizes overhead by backing up network messages, enabling efficient recovery in large supercomputers with frequent hardware failures.

Contribution

It presents a novel low-overhead fault-tolerance technique for MapReduce that relies on message backup and local storage, surpassing traditional checkpointing methods.

Findings

01

Low overhead <4% during fault-free operation

02

Recovery time is approximately the time to process 1/p of data

03

Expected communication overhead of 1/(p-1) on p processing elements

Abstract

Supercomputers getting ever larger and energy-efficient is at odds with the reliability of the used hardware. Thus, the time intervals between component failures are decreasing. Contrarily, the latencies for individual operations of coarse-grained big-data tools grow with the number of processors. To overcome the resulting scalability limit, we need to go beyond the current practice of interoperation checkpointing. We give first results on how to achieve this for the popular MapReduce framework where huge multisets are processed by user-defined mapping and reducing functions. We observe that the full state of a MapReduce algorithm is described by its network communication. We present a low-overhead technique with no additional work during fault-free execution and the negligible expected relative communication overhead of $1/ (p - 1)$ on $p$ PEs. Recovery takes approximately the time of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsCloud Computing and Resource Management · Software System Performance and Reliability · IoT and Edge/Fog Computing