Grounded Intuition of GPT-Vision's Abilities with Scientific Images

Alyssa Hwang; Andrew Head; Chris Callison-Burch

arXiv:2311.02069·cs.CL·November 6, 2023·2 cites

Grounded Intuition of GPT-Vision's Abilities with Scientific Images

Alyssa Hwang, Andrew Head, Chris Callison-Burch

PDF

Open Access 1 Repo

TL;DR

This paper introduces a qualitative framework to evaluate GPT-Vision's abilities with scientific images, revealing its sensitivities and aiding researchers in understanding its capabilities and limitations.

Contribution

It develops a grounded, example-driven qualitative evaluation method for GPT-Vision, moving beyond traditional benchmarks to better understand model behavior.

Findings

01

GPT-Vision is sensitive to prompting and counterfactual text.

02

It effectively captures relative spatial relationships in images.

03

The framework helps develop grounded intuition of model capabilities.

Abstract

GPT-Vision has impressed us on a range of vision-language tasks, but it comes with the familiar new challenge: we have little idea of its capabilities and limitations. In our study, we formalize a process that many have instinctively been trying already to develop "grounded intuition" of this new model. Inspired by the recent movement away from benchmarking in favor of example-driven qualitative evaluation, we draw upon grounded theory and thematic analysis in social science and human-computer interaction to establish a rigorous framework for qualitative evaluation in natural language processing. We use our technique to examine alt text generation for scientific figures, finding that GPT-Vision is particularly sensitive to prompting, counterfactual text in images, and relative spatial relationships. Our method and analysis aim to help researchers ramp up their own grounded intuitions of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ahwang16/grounded-intuition-gpt-vision
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Multimodal Machine Learning Applications