ANNA: Abstractive Text-to-Image Synthesis with Filtered News Captions

Aashish Anantha Ramakrishnan; Sharon X. Huang; Dongwon Lee

arXiv:2301.02160·cs.CV·August 6, 2024·1 cites

ANNA: Abstractive Text-to-Image Synthesis with Filtered News Captions

Aashish Anantha Ramakrishnan, Sharon X. Huang, Dongwon Lee

PDF

Open Access 1 Repo

TL;DR

This paper introduces ANNA, a new dataset of abstractive news captions for evaluating Text-to-Image synthesis models in news domains, highlighting challenges in understanding complex, context-rich captions.

Contribution

The paper presents ANNA, a novel dataset of abstractive news captions, and benchmarks current Text-to-Image models, revealing limitations in handling complex contextual information.

Findings

01

Transfer learning shows limited success in understanding abstractive captions.

02

Current models struggle to learn relationships between content and context.

03

ANNA dataset exposes challenges in news domain image synthesis.

Abstract

Advancements in Text-to-Image synthesis over recent years have focused more on improving the quality of generated samples using datasets with descriptive prompts. However, real-world image-caption pairs present in domains such as news data do not use simple and directly descriptive captions. With captions containing information on both the image content and underlying contextual cues, they become abstractive in nature. In this paper, we launch ANNA, an Abstractive News captioNs dAtaset extracted from online news articles in a variety of different contexts. We explore the capabilities of current Text-to-Image synthesis models to generate news domain-specific images using abstractive captions by benchmarking them on ANNA, in both standard training and transfer learning settings. The generated images are judged on the basis of contextual relevance, visual quality, and perceptual similarity…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

aashish2000/anna
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Generative Adversarial Networks and Image Synthesis · Image Retrieval and Classification Techniques

Methodsfail