Imbalanced classification: a paradigm-based review

Yang Feng; Min Zhou; Xin Tong

arXiv:2002.04592·stat.ME·July 2, 2021·Stat. Anal. Data Min.·1 cites

Imbalanced classification: a paradigm-based review

Yang Feng, Min Zhou, Xin Tong

PDF

Open Access

TL;DR

This paper reviews resampling techniques for imbalanced binary classification, analyzing their performance under different paradigms and providing guidance for practitioners based on extensive simulations and real data.

Contribution

It offers a comprehensive paradigm-based review of resampling methods, exploring their interactions with classification techniques and evaluation metrics for imbalanced data.

Findings

01

Performance varies across paradigms and techniques

02

Complex dynamics influence resampling effectiveness

03

Guidance provided for method selection

Abstract

A common issue for classification in scientific research and industry is the existence of imbalanced classes. When sample sizes of different classes are imbalanced in training data, naively implementing a classification method often leads to unsatisfactory prediction results on test data. Multiple resampling techniques have been proposed to address the class imbalance issues. Yet, there is no general guidance on when to use each technique. In this article, we provide a paradigm-based review of the common resampling techniques for binary classification under imbalanced class sizes. The paradigms we consider include the classical paradigm that minimizes the overall classification error, the cost-sensitive learning paradigm that minimizes a cost-adjusted weighted type I and type II errors, and the Neyman-Pearson paradigm that minimizes the type II error subject to a type I error…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsImbalanced Data Classification Techniques · Machine Learning and Data Classification · Financial Distress and Bankruptcy Prediction

MethodsTest