Regularized Training of Nearest Neighbor Language Models

Jean-Francois Ton; Walter Talbott; Shuangfei Zhai; Josh Susskind

arXiv:2109.08249·cs.CL·September 20, 2021·1 cites

Regularized Training of Nearest Neighbor Language Models

Jean-Francois Ton, Walter Talbott, Shuangfei Zhai, Josh Susskind

PDF

Open Access

TL;DR

This paper introduces a regularization technique during training of language models that enhances the effectiveness of post-hoc kNN retrieval, leading to improved language modeling performance especially for high-frequency words.

Contribution

It proposes a novel L2 regularization on activations during training to improve kNN-based language model performance, a concept not previously explored.

Findings

01

Regularization improves kNN classification accuracy.

02

Enhanced performance on WIKI-2 and WIKI-103 datasets.

03

Better handling of high-frequency words.

Abstract

Including memory banks in a natural language processing architecture increases model capacity by equipping it with additional data at inference time. In this paper, we build upon $k$ NN-LM \citep{khandelwal20generalization}, which uses a pre-trained language model together with an exhaustive $k$ NN search through the training data (memory bank) to achieve state-of-the-art results. We investigate whether we can improve the $k$ NN-LM performance by instead training a LM with the knowledge that we will be using a $k$ NN post-hoc. We achieved significant improvement using our method on language modeling tasks on \texttt{WIKI-2} and \texttt{WIKI-103}. The main phenomenon that we encounter is that adding a simple L2 regularization on the activations (not weights) of the model, a transformer, improves the post-hoc $k$ NN classification performance. We explore some possible reasons for this…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Multimodal Machine Learning Applications