Corpus Coverage by Top Lemmas

How much of a text can you understand knowing only the most frequent lemmas? This page shows the cumulative coverage curve: the x-axis gives the vocabulary size (top-N lemmas by frequency), and the y-axis shows what percentage of all running tokens those lemmas account for. Anchor points at N = 10, 100, 1,000 and 10,000 are marked.

Linguistic Insights

Small core vocabularies cover most of the text

In the full corpus, the 100 most frequent lemmas already cover ca. 53% of all running tokens, and the top 1,000 cover more than 77%. This concentration is high by general-language standards. The steep early rise reflects the formulaic character of song: a compact core of function words and recurring content words, amplified by the verbatim repetition of refrains, carries most of the text, leaving a long tail of words that each appear only rarely.

Coverage, corpus size, and vocabulary richness

Coverage curves are strongly shaped by corpus size, so archives are not directly comparable. At very small N the largest archives (HipHop, Chart Songs) show the highest coverage, because in a large corpus token mass concentrates on a few top function words. At larger N the relationship reverses: small single-artist archives reach very high coverage by the top 1,000 lemmas, whereas large, stylistically heterogeneous archives retain a longer tail. The shape of the curve, rather than any single coverage value, is what reflects genuine differences in vocabulary richness.

Full Corpus (all archives)

Run generate_word_coverage.R to populate this chart.

X : Top-N lemmas (log scale) ยท Y : Cumulative coverage % ยท Coverage curve ยท Anchor (N = 10 / 100 / 1k / 10k)

Lemma-based, punctuation and numerals excluded.

Top-N lemmas Coverage
10 โ€”
100 โ€”
1,000 โ€”
10,000 โ€”

These analyses are based on corpus data as of July 29, 2026.