ECOR: Explainable CLIP for Object Recognition

Ali Rasekh; Sepehr Kazemi Ranjbar; Milad Heidari; Wolfgang Nejdl

arXiv:2404.12839·cs.CV·April 22, 2024

ECOR: Explainable CLIP for Object Recognition

Ali Rasekh, Sepehr Kazemi Ranjbar, Milad Heidari, Wolfgang Nejdl

PDF

Open Access

TL;DR

This paper introduces ECOR, an explainable fine-tuning method for CLIP that enhances trustworthiness in object recognition by providing rationales without sacrificing accuracy, especially excelling in zero-shot scenarios.

Contribution

It proposes a mathematical definition of explainability for object recognition and leverages it to fine-tune CLIP, achieving state-of-the-art explainability performance.

Findings

01

State-of-the-art explainable classification results

02

Superior zero-shot performance

03

Enhanced trust in object recognition models

Abstract

Large Vision Language Models (VLMs), such as CLIP, have significantly contributed to various computer vision tasks, including object recognition and object detection. Their open vocabulary feature enhances their value. However, their black-box nature and lack of explainability in predictions make them less trustworthy in critical domains. Recently, some work has been done to force VLMs to provide reasonable rationales for object recognition, but this often comes at the expense of classification accuracy. In this paper, we first propose a mathematical definition of explainability in the object recognition task based on the joint probability distribution of categories and rationales, then leverage this definition to fine-tune CLIP in an explainable manner. Through evaluations of different datasets, our method demonstrates state-of-the-art performance in explainable classification.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMedical Imaging Techniques and Applications · Medical Image Segmentation Techniques

MethodsContrastive Language-Image Pre-training