Scaling-laws for Large Time-series Models

Thomas D. P. Edwards; James Alvey; Justin Alsing; Nam H. Nguyen,; Benjamin D. Wandelt

arXiv:2405.13867·cs.LG·January 9, 2025

Scaling-laws for Large Time-series Models

Thomas D. P. Edwards, James Alvey, Justin Alsing, Nam H. Nguyen,, Benjamin D. Wandelt

PDF

Open Access

TL;DR

This paper demonstrates that large-scale decoder-only transformer models for time series forecasting follow similar scaling laws to language models, with performance improving predictably as model size, data, and compute increase.

Contribution

It establishes for the first time power-law scaling laws for large time series transformer models across multiple parameters and datasets.

Findings

01

Scaling laws hold for time series transformers similar to language models.

02

Architectural details have minimal impact on scaling behavior.

03

Performance improves predictably with model size, data, and compute.

Abstract

Scaling laws for large language models (LLMs) have provided useful guidance in training ever larger models for predictable performance gains. Time series forecasting shares a similar sequential structure to language, and is amenable to large-scale transformer architectures. Here we show that foundational decoder-only time series transformer models exhibit analogous scaling-behavior to LLMs, with architectural details (aspect ratio and number of heads) having a minimal effect over broad ranges. We assemble a large corpus of heterogenous time series data on which to train, and establish for the first time power-law scaling with parameter count, dataset size, and training compute, spanning five orders of magnitude.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsComplex Systems and Time Series Analysis