An enriched category theory of language: from syntax to semantics

Tai-Danae Bradley; John Terilla; Yiannis Vlassopoulos

arXiv:2106.07890·math.CT·November 19, 2021

An enriched category theory of language: from syntax to semantics

Tai-Danae Bradley, John Terilla, Yiannis Vlassopoulos

PDF

Open Access

TL;DR

This paper introduces a mathematical framework using enriched category theory to connect language syntax with semantics, enabling a structured understanding of language models' probabilistic extensions.

Contribution

It develops a novel enriched categorical model that transitions from syntactic probability distributions to semantic representations via the Yoneda embedding.

Findings

01

Provides a categorical model linking syntax and semantics.

02

Enables formal reasoning about language extensions and entailment.

03

Offers a foundation for semantic analysis in language models.

Abstract

State of the art language models return a natural language text continuation from any piece of input text. This ability to generate coherent text extensions implies significant sophistication, including a knowledge of grammar and semantics. In this paper, we propose a mathematical framework for passing from probability distributions on extensions of given texts, such as the ones learned by today's large language models, to an enriched category containing semantic information. Roughly speaking, we model probability distributions on texts as a category enriched over the unit interval. Objects of this category are expressions in language, and hom objects are conditional probabilities that one expression is an extension of another. This category is syntactical -- it describes what goes with what. Then, via the Yoneda embedding, we pass to the enriched category of unit interval-valued…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Speech and dialogue systems