We Are Not Your Real Parents: Telling Causal from Confounded using MDL

David Kaltenpoth; Jilles Vreeken

arXiv:1901.06950·cs.LG·January 23, 2019·1 cites

We Are Not Your Real Parents: Telling Causal from Confounded using MDL

David Kaltenpoth, Jilles Vreeken

PDF

Open Access

TL;DR

This paper introduces CoCa, a method using MDL and latent factor modeling to distinguish causal relationships from confounding in observational data, demonstrating high accuracy and robustness.

Contribution

It develops a novel information-theoretic approach based on MDL to identify causality versus confounding, incorporating latent variable modeling with empirical validation.

Findings

01

CoCa outperforms existing methods on synthetic and real data.

02

The approach is robust even when model assumptions are violated.

03

MDL-based scores are shown to be consistent in causal inference.

Abstract

Given data over variables $(X_{1}, ..., X_{m}, Y)$ we consider the problem of finding out whether $X$ jointly causes $Y$ or whether they are all confounded by an unobserved latent variable $Z$ . To do so, we take an information-theoretic approach based on Kolmogorov complexity. In a nutshell, we follow the postulate that first encoding the true cause, and then the effects given that cause, results in a shorter description than any other encoding of the observed variables. The ideal score is not computable, and hence we have to approximate it. We propose to do so using the Minimum Description Length (MDL) principle. We compare the MDL scores under the models where $X$ causes $Y$ and where there exists a latent variables $Z$ confounding both $X$ and $Y$ and show our scores are consistent. To find potential confounders we propose using latent factor modeling, in particular, probabilistic PCA…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsBayesian Modeling and Causal Inference · Explainable Artificial Intelligence (XAI) · Computability, Logic, AI Algorithms

MethodsMinimum Description Length · Principal Components Analysis