Mitigating belief projection in explainable artificial intelligence via   Bayesian Teaching

Scott Cheng-Hsin Yang; Wai Keen Vong; Ravi B. Sojitra; Tomas Folke,; Patrick Shafto

arXiv:2102.03919·cs.AI·April 27, 2021

Mitigating belief projection in explainable artificial intelligence via Bayesian Teaching

Scott Cheng-Hsin Yang, Wai Keen Vong, Ravi B. Sojitra, Tomas Folke,, Patrick Shafto

PDF

1 Repo

TL;DR

This paper introduces Bayesian Teaching as a method to improve human understanding of AI decisions by modeling how explanations influence human reasoning, demonstrated through a binary image classification task.

Contribution

It presents a novel approach to XAI that explicitly models human reasoning with Bayesian Teaching, enhancing explanation effectiveness and interpretability.

Findings

01

Bayesian Teaching explanations improve prediction of AI judgments.

02

Sub-examples aid error detection in familiar categories.

03

Whole examples help predict AI judgments in unfamiliar cases.

Abstract

State-of-the-art deep-learning systems use decision rules that are challenging for humans to model. Explainable AI (XAI) attempts to improve human understanding but rarely accounts for how people typically reason about unfamiliar agents. We propose explicitly modeling the human explainee via Bayesian Teaching, which evaluates explanations by how much they shift explainees' inferences toward a desired goal. We assess Bayesian Teaching in a binary image classification task across a variety of contexts. Absent intervention, participants predict that the AI's classifications will match their own, but explanations generated by Bayesian Teaching improve their ability to predict the AI's judgements by moving them away from this prior belief. Bayesian Teaching further allows each case to be broken down into sub-examples (here saliency maps). These sub-examples complement whole examples by…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

CoDaS-Lab/XAI-BT-SR
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.