On Learning from Label Proportions

Felix X. Yu; Krzysztof Choromanski; Sanjiv Kumar; Tony Jebara; Shih-Fu; Chang

arXiv:1402.5902·stat.ML·February 13, 2015·41 cites

On Learning from Label Proportions

Felix X. Yu, Krzysztof Choromanski, Sanjiv Kumar, Tony Jebara, Shih-Fu, Chang

PDF

Open Access 1 Repo

TL;DR

This paper introduces a theoretical framework, Empirical Proportion Risk Minimization, for learning individual labels from group-level proportion data, providing formal guarantees and demonstrating practical feasibility through a census-based case study.

Contribution

It presents a general framework and theoretical analysis for learning from label proportions, establishing conditions under which individual labels can be accurately predicted.

Findings

01

VC bound on generalization error for bag proportions

02

Bag sample complexity is mildly sensitive to bag size

03

Good bag proportion prediction leads to accurate instance label prediction

Abstract

Learning from Label Proportions (LLP) is a learning setting, where the training data is provided in groups, or "bags", and only the proportion of each class in each bag is known. The task is to learn a model to predict the class labels of the individual instances. LLP has broad applications in political science, marketing, healthcare, and computer vision. This work answers the fundamental question, when and why LLP is possible, by introducing a general framework, Empirical Proportion Risk Minimization (EPRM). EPRM learns an instance label classifier to match the given label proportions on the training data. Our result is based on a two-step analysis. First, we provide a VC bound on the generalization error of the bag proportions. We show that the bag sample complexity is only mildly sensitive to the bag size. Second, we show that under some mild assumptions, good bag proportion…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

felixyu/pSVM
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMachine Learning and Data Classification · Imbalanced Data Classification Techniques · Machine Learning and Algorithms