Generalizing Off-Policy Learning under Sample Selection Bias

Tobias Hatt; Daniel Tschernutter; Stefan Feuerriegel

arXiv:2112.01387·stat.ML·December 3, 2021·1 cites

Generalizing Off-Policy Learning under Sample Selection Bias

Tobias Hatt, Daniel Tschernutter, Stefan Feuerriegel

PDF

Open Access

TL;DR

This paper introduces a robust policy learning framework that accounts for sample selection bias, optimizing for worst-case target population performance, and demonstrates improved generalization in simulations and clinical data.

Contribution

The paper proposes a minimax policy learning approach under sample selection bias, with an efficient algorithm and convergence guarantees for generalized policies.

Findings

01

Improved policy generalization on simulated data.

02

Enhanced performance in clinical trial data.

03

Theoretical convergence guarantees for the proposed method.

Abstract

Learning personalized decision policies that generalize to the target population is of great relevance. Since training data is often not representative of the target population, standard policy learning methods may yield policies that do not generalize target population. To address this challenge, we propose a novel framework for learning policies that generalize to the target population. For this, we characterize the difference between the training data and the target population as a sample selection bias using a selection variable. Over an uncertainty set around this selection variable, we optimize the minimax value of a policy to achieve the best worst-case policy value on the target population. In order to solve the minimax problem, we derive an efficient algorithm based on a convex-concave procedure and prove convergence for parametrized spaces of policies such as logistic…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Causal Inference Techniques · Health Systems, Economic Evaluations, Quality of Life · Advanced Bandit Algorithms Research