LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning

Xiaotian Lin; Yanlin Qi; Yizhang Zhu; Themis Palpanas; Chengliang Chai; Nan Tang; Yuyu Luo

arXiv:2505.07437·cs.LG·May 13, 2025

LEAD: Iterative Data Selection for Efficient LLM Instruction Tuning

Xiaotian Lin, Yanlin Qi, Yizhang Zhu, Themis Palpanas, Chengliang Chai, Nan Tang, Yuyu Luo

PDF

Open Access 1 Repo

TL;DR

LEAD is an efficient data selection framework for LLM instruction tuning that estimates sample utility within the training loop, greatly reducing computational costs while improving model performance.

Contribution

LEAD introduces a novel utility estimation method and a two-stage selection strategy that eliminate the need for costly model inference during data selection.

Findings

01

Outperforms state-of-the-art methods in diverse benchmarks.

02

Improves model performance by up to 10.8%.

03

Reduces training data usage by 97.5% and training time by 5-10x.

Abstract

Instruction tuning has emerged as a critical paradigm for improving the capabilities and alignment of large language models (LLMs). However, existing iterative model-aware data selection methods incur significant computational overhead, as they rely on repeatedly performing full-dataset model inference to estimate sample utility for subsequent training iterations, creating a fundamental efficiency bottleneck. In this paper, we propose LEAD, an efficient iterative data selection framework that accurately estimates sample utility entirely within the standard training loop, eliminating the need for costly additional model inference. At its core, LEAD introduces Instance-Level Dynamic Uncertainty (IDU), a theoretically grounded utility function combining instantaneous training loss, gradient-based approximation of loss changes, and exponential smoothing of historical loss signals. To…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

HKUSTDial/LEAD
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Machine Learning and Data Classification · Machine Learning and Algorithms