Decomposing and Editing Predictions by Modeling Model Computation

Harshay Shah; Andrew Ilyas; Aleksander Madry

arXiv:2404.11534·cs.LG·April 18, 2024·2 cites

Decomposing and Editing Predictions by Modeling Model Computation

Harshay Shah, Andrew Ilyas, Aleksander Madry

PDF

Open Access 1 Repo

TL;DR

This paper introduces COAR, a scalable method for decomposing model predictions into components, enabling interpretability and targeted model editing across various tasks and modalities.

Contribution

The paper presents COAR, a novel scalable algorithm for component attribution that facilitates understanding and editing of model predictions by decomposing internal computations.

Findings

01

COAR effectively estimates component attributions across models and datasets.

02

Component attributions enable diverse model editing tasks.

03

The method improves model robustness and interpretability.

Abstract

How does the internal computation of a machine learning model transform inputs into predictions? In this paper, we introduce a task called component modeling that aims to address this question. The goal of component modeling is to decompose an ML model's prediction in terms of its components -- simple functions (e.g., convolution filters, attention heads) that are the "building blocks" of model computation. We focus on a special case of this task, component attribution, where the goal is to estimate the counterfactual impact of individual components on a given prediction. We then present COAR, a scalable algorithm for estimating component attributions; we demonstrate its effectiveness across models, datasets, and modalities. Finally, we show that component attributions estimated with COAR directly enable model editing across five tasks, namely: fixing model errors, ``forgetting''…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

madrylab/modelcomponents
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsModel-Driven Software Engineering Techniques · Scientific Computing and Data Management · Simulation Techniques and Applications

MethodsFocus · Convolution