Unified Likelihood Ratio Estimation for High- to Zero-frequency N-grams

Masato Kikuchi; Kento Kawakami; Kazuho Watanabe; Mitsuo; Yoshida; Kyoji Umemura

arXiv:2110.00946·cs.CL·October 5, 2021

Unified Likelihood Ratio Estimation for High- to Zero-frequency N-grams

Masato Kikuchi, Kento Kawakami, Kazuho Watanabe, Mitsuo, Yoshida, Kyoji Umemura

PDF

TL;DR

This paper introduces a unified likelihood ratio estimation method for high- to zero-frequency N-grams in natural language processing, addressing the challenges of rare and unobserved N-grams through decomposition and regularization techniques.

Contribution

The authors propose a novel decomposition-based likelihood ratio estimation method that handles zero- and low-frequency N-grams by leveraging item unit frequencies and regularization.

Findings

01

Effective in estimating unobserved N-grams.

02

Addresses low- and zero-frequency problems.

03

Maintains dependencies between items.

Abstract

Likelihood ratios (LRs), which are commonly used for probabilistic data processing, are often estimated based on the frequency counts of individual elements obtained from samples. In natural language processing, an element can be a continuous sequence of $N$ items, called an $N$ -gram, in which each item is a word, letter, etc. In this paper, we attempt to estimate LRs based on $N$ -gram frequency information. A naive estimation approach that uses only $N$ -gram frequencies is sensitive to low-frequency (rare) $N$ -grams and not applicable to zero-frequency (unobserved) $N$ -grams; these are known as the low- and zero-frequency problems, respectively. To address these problems, we propose a method for decomposing $N$ -grams into item units and then applying their frequencies along with the original $N$ -gram frequencies. Our method can obtain the estimates of unobserved $N$ -grams by using the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.