A Two Parameters Equation for Word Rank-Frequency Relation

Chenchen Ding

arXiv:2205.00638·cs.CL·May 3, 2022

A Two Parameters Equation for Word Rank-Frequency Relation

Chenchen Ding

PDF

Open Access

TL;DR

This paper proposes a two-parameter mathematical model to accurately fit the rank-frequency relation of words in language data, improving understanding of linguistic patterns.

Contribution

It introduces a novel two-parameter equation for modeling word rank-frequency relations, extending previous models with better fit and parameter interpretability.

Findings

01

The model fits well with well-behaved linguistic data.

02

Parameters s and t are estimated from data with specific constraints.

03

The model generalizes previous rank-frequency relations.

Abstract

Let $f (\cdot)$ be the absolute frequency of words and $r$ be the rank of words in decreasing order of frequency, then the following function can fit the rank-frequency relation \[ f (r;s,t) = \left(\frac{r_{\tt max}}{r}\right)^{1-s} \left(\frac{r_{\tt max}+t \cdot r_{\tt exp}}{r+t \cdot r_{\tt exp}}\right)^{1+(1+t)s} \] where $r_{max}$ and $r_{exp}$ are the maximum and the expectation of the rank, respectively; $s > 0$ and $t > 0$ are parameters estimated from data. On well-behaved data, there should be $s < 1$ and $s \cdot t < 1$ .

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Natural Language Processing Techniques