Zipf's Law for Characters
Rank-frequency plot for the top 100 character types in the Songkorpus. The indigo curve shows the observed distribution; the dashed line shows the theoretical Zipf reference (f โ 1/r), which a perfect power-law distribution would follow exactly. Use the selector to compare the full corpus against individual archives. The underlying absolute frequencies are tabulated in Character Frequencies.
Linguistic Insights
Rank-frequency universals
Zipf's Law is a cross-linguistic distributional universal. Rank-frequency plots are also used as a corpus quality check: a strong deviation from the expected curve may indicate missing text segments or encoding artifacts in the digitised source material. Characters are counted case-insensitively: e and E are merged into a single rank, reflecting the underlying grapheme inventory rather than orthographic surface forms.
Cross-archive comparison
The distribution of the most frequent characters is highly consistent across all archives. Differences emerge mainly in inventory size and tail length, reflecting incidental variation in the source texts rather than systematic genre-specific character use.
Full Corpus (all archives)
X : Rank (top 100 characters, descending frequency) ยท Y : Absolute frequency ยท Observed ยท Zipf (f โ 1/r)
These analyses are based on corpus data as of July 29, 2026.