Text-guided Zero-Shot Object Localization

Jingjing Wang; Xinglin Piao; Zongzhi Gao; Bo Li; Yong Zhang; Baocai; Yin

arXiv:2411.11357·cs.CV·November 19, 2024

Text-guided Zero-Shot Object Localization

Jingjing Wang, Xinglin Piao, Zongzhi Gao, Bo Li, Yong Zhang, Baocai, Yin

PDF

Open Access

TL;DR

This paper introduces a zero-shot object localization framework that leverages CLIP and a novel TSSM module to identify and locate objects without labeled data, significantly improving performance.

Contribution

It presents a new zero-shot localization method combining CLIP and TSSM modules, enabling precise object localization guided by prompts without labeled training data.

Findings

01

Significant improvement in localization accuracy

02

Effective benchmark established for zero-shot localization

03

Demonstrated robustness across diverse datasets

Abstract

Object localization is a hot issue in computer vision area, which aims to identify and determine the precise location of specific objects from image or video. Most existing object localization methods heavily rely on extensive labeled data, which are costly to annotate and constrain their applicability. Therefore, we propose a new Zero-Shot Object Localization (ZSOL) framework for addressing the aforementioned challenges. In the proposed framework, we introduce the Contrastive Language Image Pre-training (CLIP) module which could integrate visual and linguistic information effectively. Furthermore, we design a Text Self-Similarity Matching (TSSM) module, which could improve the localization accuracy by enhancing the representation of text features extracted by CLIP module. Hence, the proposed framework can be guided by prompt words to identify and locate specific objects in an image in…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDomain Adaptation and Few-Shot Learning · Advanced Neural Network Applications · COVID-19 diagnosis using AI

MethodsContrastive Language-Image Pre-training