Boosting Test Performance with Importance Sampling--a Subpopulation   Perspective

Hongyu Shen; Zhizhen Zhao

arXiv:2412.13003·cs.LG·December 18, 2024

Boosting Test Performance with Importance Sampling--a Subpopulation Perspective

Hongyu Shen, Zhizhen Zhao

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper introduces importance sampling as an effective approach to address subpopulation issues in machine learning, providing a unified theoretical framework and demonstrating state-of-the-art empirical results on benchmarks.

Contribution

It offers a new systematic formulation of the subpopulation problem, clarifies connections among existing methods, and introduces a flexible estimator applicable in attribute-known and unknown scenarios.

Findings

01

Achieves state-of-the-art results on benchmark datasets.

02

Provides a theoretical connection between existing subpopulation methods.

03

Demonstrates the effectiveness of importance sampling in practical settings.

Abstract

Despite empirical risk minimization (ERM) is widely applied in the machine learning community, its performance is limited on data with spurious correlation or subpopulation that is introduced by hidden attributes. Existing literature proposed techniques to maximize group-balanced or worst-group accuracy when such correlation presents, yet, at the cost of lower average accuracy. In addition, many existing works conduct surveys on different subpopulation methods without revealing the inherent connection between these methods, which could hinder the technology advancement in this area. In this paper, we identify important sampling as a simple yet powerful tool for solving the subpopulation problem. On the theory side, we provide a new systematic formulation of the subpopulation problem and explicitly identify the assumptions that are not clearly stated in the existing works. This helps to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

skyve2012/DBA
pytorchOfficial

Videos

Boosting Test Performance with Importance Sampling--a Subpopulation Perspective· underline

Taxonomy

TopicsAdvanced Statistical Methods and Models