Beyond Fixed Length: Bucket Pre-training is All You Need

Qing Yang; Qiyao Peng; Hongtao Liu; Kai Liu; Bing Qin; Ting Liu

arXiv:2407.07495·cs.CL·June 30, 2025

Beyond Fixed Length: Bucket Pre-training is All You Need

Qing Yang, Qiyao Peng, Hongtao Liu, Kai Liu, Bing Qin, Ting Liu

PDF

Open Access

TL;DR

This paper introduces a multi-bucket data composition method for LLM pre-training that overcomes fixed-length limitations, improving efficiency and effectiveness by adaptively organizing training data based on new quality metrics.

Contribution

It proposes a novel multi-bucket data composition approach with quantitative metrics, enhancing pre-training flexibility and performance of large language models.

Findings

01

Significant improvements in pre-training efficiency.

02

Enhanced model performance on downstream tasks.

03

Effective data organization beyond fixed-length constraints.

Abstract

Large Language Models (LLMs) have demonstrated exceptional performance across various tasks, with pre-training stage serving as the cornerstone of their capabilities. However, the conventional fixed-length data composition strategy for pre-training presents several practical challenges. When using shorter sequences, documents are often truncated, potentially leading to information loss and affecting the model's ability to capture long-range dependencies. Conversely, longer sequences require concatenation of multiple documents, which can introduce noise and affect the natural document boundaries and semantic coherence as well as require substantial computational overhead. To address these challenges, we first establish three quantitative metrics for evaluating data composition quality: padding ratio, truncation ratio, and concatenation ratio. Building upon these metrics, we propose a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsShoulder Injury and Treatment