Scalable Evaluation and Improvement of Document Set Expansion via Neural   Positive-Unlabeled Learning

Alon Jacovi; Gang Niu; Yoav Goldberg; Masashi Sugiyama

arXiv:1910.13339·cs.CL·January 18, 2021

Scalable Evaluation and Improvement of Document Set Expansion via Neural Positive-Unlabeled Learning

Alon Jacovi, Gang Niu, Yoav Goldberg, Masashi Sugiyama

PDF

1 Repo

TL;DR

This paper introduces a scalable neural positive-unlabeled learning approach to enhance document set expansion, addressing challenges like class prior estimation and large-scale evaluation, with demonstrated improvements in PubMed abstract retrieval.

Contribution

It presents a novel application of neural PU learning to document retrieval, improving expansion accuracy over traditional IR methods and baselines.

Findings

01

Improved retrieval accuracy on PubMed abstracts.

02

Effective handling of class prior and data imbalance.

03

Scalable evaluation methodology for large datasets.

Abstract

We consider the situation in which a user has collected a small set of documents on a cohesive topic, and they want to retrieve additional documents on this topic from a large collection. Information Retrieval (IR) solutions treat the document set as a query, and look for similar documents in the collection. We propose to extend the IR approach by treating the problem as an instance of positive-unlabeled (PU) learning -- i.e., learning binary classifiers from only positive and unlabeled data, where the positive data corresponds to the query documents, and the unlabeled data is the results returned by the IR engine. Utilizing PU learning for text with big neural networks is a largely unexplored field. We discuss various challenges in applying PU learning to the setting, including an unknown class prior, extremely imbalanced data and large-scale accurate evaluation of models, and we…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

sayaendo/document-set-expansion-pu
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.