Zipf's Law for Words

Rank-frequency plot for the top 1,000 word tokens in the Songkorpus. The green curve shows the observed distribution; the dashed line shows the theoretical Zipf reference (f โˆ 1/r). The word-level distribution falls steeply in the first few dozen ranks โ€” high-frequency function words dominate โ€” then flattens into a long tail of less frequent content words. Use the selector to compare the full corpus against individual archives. The underlying absolute frequencies are tabulated in Word Frequencies.

Linguistic Insights

Rank-frequency universals

Zipf's Law was originally formulated for word tokens and remains the canonical rank-frequency universal in linguistics, with its near-power-law shape holding across typologically diverse languages. A pure Zipf curve, however, almost never fits the very top ranks: the Zipf-Mandelbrot variant, which adds parameters to the head of the distribution, captures the highest-frequency words far better and accounts for the deviation visible in the first few dozen ranks. Note that this view counts word forms case-sensitively, so ich/Ich and und/Und occupy separate ranks; merging them would redistribute token mass and steepen the head of the curve.

The refrain effect

Unlike prose, song lyrics are structurally repetitive: choruses recur verbatim, often several times per song. This inflates the head of the rank-frequency distribution. Comparisons should account for this: genres with longer or more frequently repeated refrains will show a more concentrated head independently of any difference in underlying vocabulary, making the refrain a confound that is also, in itself, a worthwhile object of study.

Full Corpus (all archives)

Run generate_word_zipf.R to populate this chart.

X : Rank (top 1,000 word tokens, descending frequency) ยท Y : Absolute frequency ยท Observed ยท Zipf (f โˆ 1/r)

These analyses are based on corpus data as of July 29, 2026.