Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry

Kyle Elliott Mathewson

arXiv:2603.02258·cs.CL·March 4, 2026

Universal Conceptual Structure in Neural Translation: Probing NLLB-200's Multilingual Geometry

Kyle Elliott Mathewson

PDF

Open Access

TL;DR

This paper investigates whether neural machine translation models learn universal conceptual representations across languages by analyzing the geometry of their embeddings, revealing correlations with linguistic and cognitive structures.

Contribution

The study provides empirical evidence that NLLB-200 encodes language-universal conceptual structures, bridging NLP interpretability with cognitive science theories.

Findings

01

Embedding distances correlate with phylogenetic language distances

02

Model internalizes universal conceptual associations

03

Cross-lingual semantic offsets are highly consistent

Abstract

Do neural machine translation models learn language-universal conceptual representations, or do they merely cluster languages by surface similarity? We investigate this question by probing the representation geometry of Meta's NLLB-200, a 200-language encoder-decoder Transformer, through six experiments that bridge NLP interpretability with cognitive science theories of multilingual lexical organization. Using the Swadesh core vocabulary list embedded across 135 languages, we find that the model's embedding distances significantly correlate with phylogenetic distances from the Automated Similarity Judgment Program ( $ρ = 0.13$ , $p = 0.020$ ), demonstrating that NLLB-200 has implicitly learned the genealogical structure of human languages. We show that frequently colexified concept pairs from the CLICS database exhibit significantly higher embedding similarity than non-colexified pairs…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNeurobiology of Language and Bilingualism · Natural Language Processing Techniques · Action Observation and Synchronization