CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning

Guoliang He; Eiko Yoneki

arXiv:2501.08071·cs.AR·January 15, 2025

CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning

Guoliang He, Eiko Yoneki

PDF

Open Access 1 Repo 1 Models

TL;DR

CuAsmRL employs deep reinforcement learning to automatically optimize GPU SASS schedules, outperforming manual tuning and enhancing CUDA kernel performance by up to 26%.

Contribution

This work introduces a novel RL-based method to automate GPU assembly scheduling, mimicking expert manual optimization within compiler frameworks.

Findings

01

Up to 26% performance improvement over existing CUDA kernels.

02

Average 9% performance gain across tested kernels.

03

Demonstrates RL can learn effective GPU scheduling strategies.

Abstract

Large language models (LLMs) are remarked by their substantial computational requirements. To mitigate the cost, researchers develop specialized CUDA kernels, which often fuse several tensor operations to maximize the utilization of GPUs as much as possible. However, those specialized kernels may still leave performance on the table as CUDA assembly experts show that manual optimization of GPU SASS schedules can lead to better performance, and trial-and-error is largely employed to manually find the best GPU SASS schedules. In this work, we employ an automatic approach to optimize GPU SASS schedules, which thus can be integrated into existing compiler frameworks. The key to automatic optimization is training an RL agent to mimic how human experts perform manual scheduling. To this end, we formulate an assembly game, where RL agents can play to find the best GPU SASS schedules. The…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

hgl71964/cuasmrl
pytorchOfficial

Models

🤗
JayLuci4/chronos-poc
model

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDistributed and Parallel Computing Systems · Parallel Computing and Optimization Techniques · Medical Image Segmentation Techniques