Formalizing BPE Tokenization

Martin Berglund (Ume{\aa} University); Brink van der Merwe; (Stellenbosch University)

arXiv:2309.08715·cs.FL·September 19, 2023·NCMA

Formalizing BPE Tokenization

Martin Berglund (Ume{\aa} University), Brink van der Merwe, (Stellenbosch University)

PDF

TL;DR

This paper formalizes byte pair encoding tokenization used in NLP, analyzing the semantics of popular tokenizers and exploring incremental, memory-efficient tokenization methods.

Contribution

It provides a formal definition of BPE tokenization, compares SentencePiece and HuggingFace tokenizers, and introduces methods for incremental, memory-efficient tokenization.

Findings

01

Formal semantics of SentencePiece and HuggingFace tokenizers

02

Relationship between different tokenization rule constructions

03

Proposed incremental, constant-memory tokenization method

Abstract

In this paper, we formalize practical byte pair encoding tokenization as it is used in large language models and other NLP systems, in particular we formally define and investigate the semantics of the SentencePiece and HuggingFace tokenizers, in particular how they relate to each other, depending on how the tokenization rules are constructed. Beyond this we consider how tokenization can be performed in an incremental fashion, as well as doing it left-to-right using an amount of memory constant in the length of the string, enabling e.g. using a finite state string-to-string transducer.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.