$k$Folden: $k$-Fold Ensemble for Out-Of-Distribution Detection

Xiaoya Li; Jiwei Li; Xiaofei Sun; Chun Fan; Tianwei Zhang; Fei Wu,; Yuxian Meng; Jun Zhang

arXiv:2108.12731·cs.CL·November 9, 2021·5 cites

$k$Folden: $k$-Fold Ensemble for Out-Of-Distribution Detection

Xiaoya Li, Jiwei Li, Xiaofei Sun, Chun Fan, Tianwei Zhang, Fei Wu,, Yuxian Meng, Jun Zhang

PDF

Open Access 1 Repo

TL;DR

kFolden is a novel ensemble framework for out-of-distribution detection in NLP that trains sub-models on masked categories to improve OOD detection without external data, outperforming existing methods.

Contribution

The paper introduces kFolden, a simple ensemble approach that simulates OOD scenarios during training, enhancing detection performance without external datasets.

Findings

01

kFolden outperforms existing OOD detection methods in benchmarks.

02

It maintains high in-domain classification accuracy.

03

The framework is effective across multiple text classification datasets.

Abstract

Out-of-Distribution (OOD) detection is an important problem in natural language processing (NLP). In this work, we propose a simple yet effective framework $k$ Folden, which mimics the behaviors of OOD detection during training without the use of any external data. For a task with $k$ training labels, $k$ Folden induces $k$ sub-models, each of which is trained on a subset with $k - 1$ categories with the left category masked unknown to the sub-model. Exposing an unknown label to the sub-model during training, the model is encouraged to learn to equally attribute the probability to the seen $k - 1$ labels for the unknown label, enabling this framework to simultaneously resolve in- and out-distribution examples in a natural way via OOD simulations. Taking text classification as an archetype, we develop benchmarks for OOD detection using existing text classification datasets. By conducting…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ShannonAI/kfolden-ood-detection
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Speech Recognition and Synthesis