SOTAVerified

A Crucial Parameter for Rank-Frequency Relation in Natural Languages

2024-02-01Unverified0· sign in to hype

Chenchen Ding

Unverified — Be the first to reproduce this paper.

Reproduce

Abstract

f r^- (r+)^- has been empirically shown more precise than a na\"ive power law f r^- to model the rank-frequency (r-f) relation of words in natural languages. This work shows that the only crucial parameter in the formulation is , which depicts the resistance to vocabulary growth on a corpus. A method of parameter estimation by searching an optimal is proposed, where a ``zeroth word'' is introduced technically for the calculation. The formulation and parameters are further discussed with several case studies.

Tasks

Reproductions