PerceptionCLIP: Visual Classification by Inferring and Conditioning on   Contexts

Bang An; Sicheng Zhu; Michael-Andrei Panaitescu-Liess; Chaithanya; Kumar Mummadi; Furong Huang

arXiv:2308.01313·cs.CV·March 19, 2024

PerceptionCLIP: Visual Classification by Inferring and Conditioning on Contexts

Bang An, Sicheng Zhu, Michael-Andrei Panaitescu-Liess, Chaithanya, Kumar Mummadi, Furong Huang

PDF

Open Access 1 Repo

TL;DR

PerceptionCLIP enhances zero-shot image classification by inferring and conditioning on contextual attributes, inspired by human perception, leading to improved robustness and generalization without additional training.

Contribution

It introduces a training-free, two-step method that leverages CLIP's ability to infer context, improving zero-shot classification performance.

Findings

01

Improves zero-shot classification accuracy

02

Enhances robustness to spurious features

03

Achieves better generalization and group robustness

Abstract

Vision-language models like CLIP are widely used in zero-shot image classification due to their ability to understand various visual concepts and natural language descriptions. However, how to fully leverage CLIP's unprecedented human-like understanding capabilities to achieve better performance is still an open question. This paper draws inspiration from the human visual perception process: when classifying an object, humans first infer contextual attributes (e.g., background and orientation) which help separate the foreground object from the background, and then classify the object based on this information. Inspired by it, we observe that providing CLIP with contextual attributes improves zero-shot image classification and mitigates reliance on spurious features. We also observe that CLIP itself can reasonably infer the attributes from an image. With these observations, we propose a…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

umd-huang-lab/perceptionclip
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDomain Adaptation and Few-Shot Learning · Multimodal Machine Learning Applications · Advanced Image and Video Retrieval Techniques

MethodsContrastive Language-Image Pre-training