Evaluating authenticity and quality of image captions via sentiment and   semantic analyses

Aleksei Krotov; Alison Tebo; Dylan K. Picart; Aaron Dean Algave

arXiv:2409.09560·cs.CV·September 17, 2024

Evaluating authenticity and quality of image captions via sentiment and semantic analyses

Aleksei Krotov, Alison Tebo, Dylan K. Picart, Aaron Dean Algave

PDF

Open Access

TL;DR

This paper introduces a method to evaluate the quality of image captions by analyzing sentiment and semantic richness, revealing insights into caption diversity and sentiment influence across a large dataset.

Contribution

It presents a novel evaluation approach using pre-trained models to assess sentiment and semantic variability in image captions, aiding quality control of crowd-sourced data.

Findings

01

Most captions were neutral with low semantic variability.

02

Approximately 6% of captions showed strong sentiment influenced by object categories.

03

Generated captions had minimal strong sentiment, unaffected by object categories.

Abstract

The growth of deep learning (DL) relies heavily on huge amounts of labelled data for tasks such as natural language processing and computer vision. Specifically, in image-to-text or image-to-image pipelines, opinion (sentiment) may be inadvertently learned by a model from human-generated image captions. Additionally, learning may be affected by the variety and diversity of the provided captions. While labelling large datasets has largely relied on crowd-sourcing or data-worker pools, evaluating the quality of such training data is crucial. This study proposes an evaluation method focused on sentiment and semantic richness. That method was applied to the COCO-MS dataset, comprising approximately 150K images with segmented objects and corresponding crowd-sourced captions. We employed pre-trained models (Twitter-RoBERTa-base and BERT-base) to extract sentiment scores and variability of…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSubtitles and Audiovisual Media · Language, Metaphor, and Cognition · Video Analysis and Summarization