Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies

Giuseppe Samo; Paola Merlo

arXiv:2603.15295·cs.CL·March 17, 2026

Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies

Giuseppe Samo, Paola Merlo

PDF

Open Access

TL;DR

This paper introduces curated, multilingual datasets based on Blackbird Language Matrices to evaluate large language models' understanding of verb alternations through controlled linguistic puzzles.

Contribution

It presents novel, systematic datasets and data augmentation strategies for probing cross-sentence verb alternation knowledge in multiple languages.

Findings

01

Baseline results show datasets are effective diagnostic tools.

02

Datasets cover multiple languages and verb alternation types.

03

Data augmentation improves dataset diversity.

Abstract

Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns, such as verb alternations, remains underexplored. In this work, we present curated paradigm-based datasets for four languages, designed to probe systematic cross-sentence knowledge of verb alternations (change-of-state and object-drop constructions in English, German and Italian, and Hebrew binyanim). The datasets comprise thousands of the Blackbird Language Matrices (BLMs) problems. The BLM task -- an RPM/ARC-like task devised specifically for language -- is a controlled linguistic puzzle where models must select the sentence that completes a pattern according to syntactic and semantic rules. We introduce three types of templates varying in complexity and apply linguistically-informed data augmentation…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Language and cultural evolution · Text Readability and Simplification