Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models

Srishti Yadav; Zhi Zhang; Daniel Hershcovich; Ekaterina Shutova

arXiv:2502.14906·cs.CL·February 24, 2025

Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models

Srishti Yadav, Zhi Zhang, Daniel Hershcovich, Ekaterina Shutova

PDF

1 Video

TL;DR

This paper evaluates how large vision-language models (VLMs) reflect cultural values, revealing that their alignment with cultural sensitivities varies with context, highlighting challenges and potential for improvement.

Contribution

It is the first comprehensive study assessing cultural value sensitivity in multimodal models, comparing their performance to language-only models across different scales.

Findings

01

VLMs exhibit cultural value sensitivity similar to LLMs.

02

Performance in aligning with cultural values is highly context-dependent.

03

Using images can enhance understanding of cultural values but introduces variability.

Abstract

Investigating value alignment in Large Language Models (LLMs) based on cultural context has become a critical area of research. However, similar biases have not been extensively explored in large vision-language models (VLMs). As the scale of multimodal models continues to grow, it becomes increasingly important to assess whether images can serve as reliable proxies for culture and how these values are embedded through the integration of both visual and textual data. In this paper, we conduct a thorough evaluation of multimodal model at different scales, focusing on their alignment with cultural values. Our findings reveal that, much like LLMs, VLMs exhibit sensitivity to cultural values, but their performance in aligning with these values is highly context-dependent. While VLMs show potential in improving value understanding through the use of images, this alignment varies…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

Beyond Words: Exploring Cultural Value Sensitivity in Multimodal Models· underline