Training Priors Predict Text-To-Image Model Performance

Charles Lovering; Ellie Pavlick

arXiv:2306.01755·cs.CV·October 26, 2023·1 cites

Training Priors Predict Text-To-Image Model Performance

Charles Lovering, Ellie Pavlick

PDF

Open Access

TL;DR

This paper investigates how training data frequency influences the ability of text-to-image models to generate correctly related subject-verb-object triads, revealing biases towards seen relations and challenging assumptions about compositional generalization.

Contribution

It demonstrates that training priors significantly affect relation generation in text-to-image models, providing evidence that models rely on seen relations rather than true compositionality.

Findings

01

Higher training frequency improves correct relation generation

02

Models struggle with flipped or less frequent relations

03

Training data biases influence model outputs

Abstract

Text-to-image models can often generate some relations, i.e., "astronaut riding horse", but fail to generate other relations composed of the same basic parts, i.e., "horse riding astronaut". These failures are often taken as evidence that models rely on training priors rather than constructing novel images compositionally. This paper tests this intuition on the stablediffusion 2.1 text-to-image model. By looking at the subject-verb-object (SVO) triads that underlie these prompts (e.g., "astronaut", "ride", "horse"), we find that the more often an SVO triad appears in the training data, the better the model can generate an image aligned with that triad. Here, by aligned we mean that each of the terms appears in the generated image in the proper relation to each other. Surprisingly, this increased frequency also diminishes how well the model can generate an image aligned with the flipped…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Multimodal Machine Learning Applications · Domain Adaptation and Few-Shot Learning

Methodsfail