Zipf's Law for Words
Rank-frequency plot for the top 1,000 word tokens in the Songkorpus. The green curve shows the observed distribution; the dashed line shows the theoretical Zipf reference (f โ 1/r). The word-level distribution falls steeply in the first few dozen ranks โ high-frequency function words dominate โ then flattens into a long tail of less frequent content words. Use the selector to compare the full corpus against individual archives. The underlying absolute frequencies are tabulated in Word Frequencies.
Linguistic Insights
Rank-frequency universals
Zipf's Law was originally formulated for word tokens and remains the canonical rank-frequency universal in linguistics, with its near-power-law shape holding across typologically diverse languages. A pure Zipf curve, however, almost never fits the very top ranks: the Zipf-Mandelbrot variant, which adds parameters to the head of the distribution, captures the highest-frequency words far better and accounts for the deviation visible in the first few dozen ranks. Note that this view counts word forms case-sensitively, so ich/Ich and und/Und occupy separate ranks; merging them would redistribute token mass and steepen the head of the curve.
The refrain effect
Unlike prose, song lyrics are structurally repetitive: choruses recur verbatim, often several times per song. This inflates the head of the rank-frequency distribution. Comparisons should account for this: genres with longer or more frequently repeated refrains will show a more concentrated head independently of any difference in underlying vocabulary, making the refrain a confound that is also, in itself, a worthwhile object of study.
Full Corpus (all archives)
X : Rank (top 1,000 word tokens, descending frequency) ยท Y : Absolute frequency ยท Observed ยท Zipf (f โ 1/r)
These analyses are based on corpus data as of July 29, 2026.