Learning From Missing Data Using Selection Bias in Movie Recommendation

Claire Vernade (LTCI); Olivier Capp\'e (LTCI)

arXiv:1509.09130·stat.ML·October 1, 2015

Learning From Missing Data Using Selection Bias in Movie Recommendation

Claire Vernade (LTCI), Olivier Capp\'e (LTCI)

PDF

TL;DR

This paper investigates how selection bias affects movie recommendation data and introduces a variational method to leverage this bias, enhancing rating estimation and recommendation reliability.

Contribution

It provides statistical evidence of selection bias in recommendation datasets and proposes a novel variational approach to exploit this bias for better rating predictions.

Findings

01

Selection bias is significant in movie recommendation data.

02

The proposed method improves rating estimation accuracy.

03

Enhanced recommendation reliability demonstrated with neighborhood-based filtering.

Abstract

Recommending items to users is a challenging task due to the large amount of missing information. In many cases, the data solely consist of ratings or tags voluntarily contributed by each user on a very limited subset of the available items, so that most of the data of potential interest is actually missing. Current approaches to recommendation usually assume that the unobserved data is missing at random. In this contribution, we provide statistical evidence that existing movie recommendation datasets reveal a significant positive association between the rating of items and the propensity to select these items. We propose a computationally efficient variational approach that makes it possible to exploit this selection bias so as to improve the estimation of ratings from small populations of users. Results obtained with this approach applied to neighborhood-based collaborative filtering…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.