Active WeaSuL: Improving Weak Supervision with Active Learning

Samantha Biegel; Rafah El-Khatib; Luiz Otavio Vilas Boas Oliveira; Max; Baak; Nanne Aben

arXiv:2104.14847·cs.LG·May 3, 2021

Active WeaSuL: Improving Weak Supervision with Active Learning

Samantha Biegel, Rafah El-Khatib, Luiz Otavio Vilas Boas Oliveira, Max, Baak, Nanne Aben

PDF

Open Access 1 Repo

TL;DR

Active WeaSuL combines weak supervision with active learning to efficiently improve label quality using minimal expert-labeled data, especially useful when labeling resources are limited.

Contribution

It introduces a novel method that integrates active learning into weak supervision, enhancing label accuracy with fewer labeled examples.

Findings

01

Outperforms weak supervision and active learning alone with limited labels

02

Effective with as few as 60 labeled data points

03

Improves probabilistic label estimation through expert feedback

Abstract

The availability of labelled data is one of the main limitations in machine learning. We can alleviate this using weak supervision: a framework that uses expert-defined rules $λ$ to estimate probabilistic labels $p (y ∣ λ)$ for the entire data set. These rules, however, are dependent on what experts know about the problem, and hence may be inaccurate or may fail to capture important parts of the problem-space. To mitigate this, we propose Active WeaSuL: an approach that incorporates active learning into weak supervision. In Active WeaSuL, experts do not only define rules, but they also iteratively provide the true label for a small set of points where the weak supervision model is most likely to be mistaken, which are then used to better estimate the probabilistic labels. In this way, the weak labels provide a warm start, which active learning then…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

SamanthaBiegel/ActiveWeaSuL
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMachine Learning and Algorithms · Machine Learning and Data Classification · Imbalanced Data Classification Techniques