Zipf's Law for Characters

Rank-frequency plot for the top 100 character types in the Songkorpus. The indigo curve shows the observed distribution; the dashed line shows the theoretical Zipf reference (f โˆ 1/r), which a perfect power-law distribution would follow exactly. Use the selector to compare the full corpus against individual archives. The underlying absolute frequencies are tabulated in Character Frequencies.

Linguistic Insights

Rank-frequency universals

Zipf's Law is a cross-linguistic distributional universal. Rank-frequency plots are also used as a corpus quality check: a strong deviation from the expected curve may indicate missing text segments or encoding artifacts in the digitised source material. Characters are counted case-insensitively: e and E are merged into a single rank, reflecting the underlying grapheme inventory rather than orthographic surface forms.

Cross-archive comparison

The distribution of the most frequent characters is highly consistent across all archives. Differences emerge mainly in inventory size and tail length, reflecting incidental variation in the source texts rather than systematic genre-specific character use.

Full Corpus (all archives)

Run generate_char_zipf.R to populate this chart.

X : Rank (top 100 characters, descending frequency) ยท Y : Absolute frequency ยท Observed ยท Zipf (f โˆ 1/r)

These analyses are based on corpus data as of July 29, 2026.