How much data do you need? Part 2: Predicting DL class specific training   dataset sizes

Thomas M\"uhlenst\"adt; Jelena Frtunikj

arXiv:2403.06311·cs.LG·March 12, 2024·1 cites

How much data do you need? Part 2: Predicting DL class specific training dataset sizes

Thomas M\"uhlenst\"adt, Jelena Frtunikj

PDF

Open Access

TL;DR

This paper introduces a method to predict the performance of classification models based on the distribution of training examples across classes, using models like powerlaw curves and applying it to datasets like CIFAR10 and EMNIST.

Contribution

It proposes a novel algorithm for estimating class-specific dataset sizes needed for optimal model performance, extending traditional models with a new combinatorial approach.

Findings

01

Effective prediction of class-specific dataset sizes.

02

Application to CIFAR10 and EMNIST datasets.

03

Model fits well with experimental data.

Abstract

This paper targets the question of predicting machine learning classification model performance, when taking into account the number of training examples per class and not just the overall number of training examples. This leads to the a combinatorial question, which combinations of number of training examples per class should be considered, given a fixed overall training dataset size. In order to solve this question, an algorithm is suggested which is motivated from special cases of space filling design of experiments. The resulting data are modeled using models like powerlaw curves and similar models, extended like generalized linear models i.e. by replacing the overall training dataset size by a parametrized linear combination of the number of training examples per label class. The proposed algorithm has been applied on the CIFAR10 and the EMNIST datasets.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsEducational Assessment and Pedagogy · Online Learning and Analytics · Advanced Data Processing Techniques