Scalable System Scheduling for HPC and Big Data

Albert Reuther; Chansup Byun; William Arcand; David Bestor; Bill; Bergeron; Matthew Hubbell; Michael Jones; Peter Michaleas; Andrew Prout,; Antonio Rosa; Jeremy Kepner

arXiv:1705.03102·cs.DC·March 6, 2018

Scalable System Scheduling for HPC and Big Data

Albert Reuther, Chansup Byun, William Arcand, David Bestor, Bill, Bergeron, Matthew Hubbell, Michael Jones, Peter Michaleas, Andrew Prout,, Antonio Rosa, Jeremy Kepner

PDF

TL;DR

This paper analyzes the performance of job schedulers in HPC and big data systems, develops a theoretical model for scheduler latency, and demonstrates that multilevel scheduling significantly improves short workload utilization.

Contribution

It introduces a detailed feature analysis and a theoretical latency model for schedulers, and shows how multilevel schedulers enhance short workload efficiency.

Findings

01

Scheduler performance is characterized by marginal latency and a nonlinear exponent.

02

System utilization drops below 10% for short computations without multilevel scheduling.

03

Multilevel schedulers can increase short workload utilization to over 90%.

Abstract

In the rapidly expanding field of parallel processing, job schedulers are the "operating systems" of modern big data architectures and supercomputing systems. Job schedulers allocate computing resources and control the execution of processes on those resources. Historically, job schedulers were the domain of supercomputers, and job schedulers were designed to run massive, long-running computations over days and weeks. More recently, big data workloads have created a need for a new class of computations consisting of many short computations taking seconds or minutes that process enormous quantities of data. For both supercomputers and big data systems, the efficiency of the job scheduler represents a fundamental limit on the efficiency of the system. Detailed measurement and modeling of the performance of schedulers are critical for maximizing the performance of a large-scale computing…

Figures12

Click any figure to enlarge with its caption.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.