ScaleFold: Reducing AlphaFold Initial Training Time to 10 Hours

Feiwen Zhu; Arkadiusz Nowaczynski; Rundong Li; Jie Xin; Yifei Song,; Michal Marcinkiewicz; Sukru Burc Eryilmaz; Jun Yang; Michael Andersch

arXiv:2404.11068·cs.LG·April 18, 2024·1 cites

ScaleFold: Reducing AlphaFold Initial Training Time to 10 Hours

Feiwen Zhu, Arkadiusz Nowaczynski, Rundong Li, Jie Xin, Yifei Song,, Michal Marcinkiewicz, Sukru Burc Eryilmaz, Jun Yang, Michael Andersch

PDF

Open Access

TL;DR

ScaleFold is a training method that significantly accelerates AlphaFold's training process, reducing pretraining time from seven days to just 10 hours by optimizing communication and computation overheads.

Contribution

The paper introduces ScaleFold, a systematic training approach that enables efficient scaling of AlphaFold training to hundreds of GPUs, achieving substantial speedups over prior methods.

Findings

01

ScaleFold trained AlphaFold in 10 hours from scratch.

02

Achieved over 6x speedup in benchmark tests.

03

Successfully scaled training to 2080 GPUs.

Abstract

AlphaFold2 has been hailed as a breakthrough in protein folding. It can rapidly predict protein structures with lab-grade accuracy. However, its implementation does not include the necessary training code. OpenFold is the first trainable public reimplementation of AlphaFold. AlphaFold training procedure is prohibitively time-consuming, and gets diminishing benefits from scaling to more compute resources. In this work, we conducted a comprehensive analysis on the AlphaFold training procedure based on Openfold, identified that inefficient communications and overhead-dominated computations were the key factors that prevented the AlphaFold training from effective scaling. We introduced ScaleFold, a systematic training method that incorporated optimizations specifically for these factors. ScaleFold successfully scaled the AlphaFold training to 2080 NVIDIA H100 GPUs with high resource…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsParallel Computing and Optimization Techniques

MethodsAlphaFold