Learning Visual Grounding from Generative Vision and Language Model

Shijie Wang; Dahun Kim; Ali Taalimi; Chen Sun; Weicheng Kuo

arXiv:2407.14563·cs.CV·July 23, 2024

Learning Visual Grounding from Generative Vision and Language Model

Shijie Wang, Dahun Kim, Ali Taalimi, Chen Sun, Weicheng Kuo

PDF

Open Access 1 Models

TL;DR

This paper demonstrates that generative vision-language models can be prompted to generate large-scale, high-quality visual grounding datasets, significantly improving zero-shot performance on referring expression tasks.

Contribution

The authors introduce a method to leverage generative VLMs for creating extensive visual grounding datasets with purely model-generated queries, enhancing zero-shot transfer capabilities.

Findings

01

Constructed a dataset with 500K images and 16M referring expressions.

02

Achieved state-of-the-art zero-shot results on RefCOCO benchmarks.

03

Showed that generative VLMs can effectively produce high-quality grounding data.

Abstract

Visual grounding tasks aim to localize image regions based on natural language references. In this work, we explore whether generative VLMs predominantly trained on image-text data could be leveraged to scale up the text annotation of visual grounding data. We find that grounding knowledge already exists in generative VLM and can be elicited by proper prompting. We thus prompt a VLM to generate object-level descriptions by feeding it object regions from existing object detection datasets. We further propose attribute modeling to explicitly capture the important object attributes, and spatial relation modeling to capture inter-object relationship, both of which are common linguistic pattern in referring expression. Our constructed dataset (500K images, 1M objects, 16M referring expressions) is one of the largest grounding datasets to date, and the first grounding dataset with purely…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

🤗
linhuixiao/Awesome-Visual-Grounding
model· ♡ 1
♡ 1

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications