Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties

Jannis Vamvas; Ignacio P\'erez Prat; Angela Heldstab; Dominic P. Fischer; Sina Ahmadi; and Rico Sennrich

arXiv:2603.25489·cs.CL·March 27, 2026

Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties

Jannis Vamvas, Ignacio P\'erez Prat, Angela Heldstab, Dominic P. Fischer, Sina Ahmadi, and Rico Sennrich

PDF

Open Access 1 Models 1 Datasets

TL;DR

This paper investigates how translation asymmetry in large language models affects data augmentation for low-resource Romansh language varieties, revealing that aligning augmentation direction with resource gradients improves translation quality.

Contribution

It demonstrates that resource-gradient-aligned data augmentation outperforms standard methods, achieving fluent translations in Romansh varieties and surpassing existing models like Gemini 3 Pro.

Findings

01

Resource-gradient-aligned augmentation improves BLEU scores by 23 in Romansh varieties.

02

The approach yields the first fluent translations for individual Romansh varieties.

03

LLMs tend to confuse the 6 Romansh language varieties, affecting translation quality.

Abstract

Recent strategies for low-resource machine translation rely on LLMs to generate synthetic data from higher-resource languages. We find that this method fails for Romansh, because LLMs tend to confuse its 6 distinct language varieties. Our experiments show that instead, the direction of data augmentation should be aligned with the resource gradient between source and target language. This approach surpasses Gemini 3 Pro in the lowest-resource variety of Romansh by 23 BLEU. A human evaluation confirms that our experiments yield the first model that generates fluent translations in the individual Romansh varieties.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Models

🤗
ZurichNLP/romansh-nllb-1.3b-ct2
model· 51 dl
51 dl

Datasets

ZurichNLP/romansh-mt-evaluation
dataset· 63 dl
63 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Digital Humanities and Scholarship · Topic Modeling