Improving Distribution Alignment with Diversity-based Sampling

Andrea Napoli; Paul White

arXiv:2410.04235·cs.LG·October 8, 2024

Improving Distribution Alignment with Diversity-based Sampling

Andrea Napoli, Paul White

PDF

Open Access

TL;DR

This paper introduces diversity-based sampling methods to improve distribution alignment in machine learning, reducing noise in discrepancy estimates and enhancing model generalization under domain shifts.

Contribution

It proposes two diversity-based sampling techniques, k-DPP and k-means++, as drop-in replacements for random sampling to improve distribution alignment.

Findings

01

More representative minibatches improve distribution estimates

02

Reduced discrepancy estimation error with smaller samples

03

Enhanced out-of-distribution accuracy across methods

Abstract

Domain shifts are ubiquitous in machine learning, and can substantially degrade a model's performance when deployed to real-world data. To address this, distribution alignment methods aim to learn feature representations which are invariant across domains, by minimising the discrepancy between the distributions. However, the discrepancy estimates can be extremely noisy when training via stochastic gradient descent (SGD), and shifts in the relative proportions of different subgroups can lead to domain misalignments; these can both stifle the benefits of the method. This paper proposes to improve these estimates by inducing diversity in each sampled minibatch. This simultaneously balances the data and reduces the variance of the gradients, thereby enhancing the model's generalisation ability. We describe two options for diversity-based data samplers, based on the k-determinantal point…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsCensus and Population Estimation