A Large-Scale Multilingual Study of Visual Constraints on Linguistic Selection of Descriptions
Uri Berger, Lea Frermann, Gabriel Stanovsky, Omri Abend

TL;DR
This large-scale multilingual study investigates how visual context influences linguistic choices in image captions across four languages, revealing consistent patterns and contributing new methods for analyzing visual-linguistic relationships.
Contribution
The paper introduces a novel method leveraging existing image caption corpora to study visual constraints on language across multiple languages and properties.
Findings
Visual context constrains linguistic properties across languages.
Patterns in numeral usage are influenced by visual conditions.
Methodology extends existing research in cognitive linguistics.
Abstract
We present a large, multilingual study into how vision constrains linguistic choice, covering four languages and five linguistic properties, such as verb transitivity or use of numerals. We propose a novel method that leverages existing corpora of images with captions written by native speakers, and apply it to nine corpora, comprising 600k images and 3M captions. We study the relation between visual input and linguistic choices by training classifiers to predict the probability of expressing a property from raw images, and find evidence supporting the claim that linguistic properties are constrained by visual context across languages. We complement this investigation with a corpus study, taking the test case of numerals. Specifically, we use existing annotations (number or type of objects) to investigate the effect of different visual conditions on the use of numeral expressions in…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsLanguage, Metaphor, and Cognition · Subtitles and Audiovisual Media · Categorization, perception, and language
MethodsTest
