Type-Token Ratio (STTR, MATTR)

Mean lexical diversity per year, measured by two length-corrected type-token ratios. Only songs with at least 100 tokens are included. Select an archive below to narrow the view. For a sequence-based alternative not penalised by text length, see MTLD β†’

Linguistic Insights

Rap sits far below every other archive

The HipHop Songs archive has a mean STTR around 0.20 β€” roughly half that of the most diverse archives (GDR, NDW). This likely reflects a genuine genre property: rap leans heavily on repetition. It is not the whole story, though β€” a sequence-based measure reverses the ranking (see MTLD β†’). Note too that the corpus is not balanced over time, so yearly comparisons mix vocabulary change with changes in what was recorded.

STTR and MATTR agree almost perfectly

Both measures correct for length with a 100-token window β€” STTR averages non-overlapping segments, MATTR slides the window β€” and in practice they barely diverge, rarely by more than 0.01, moving together year on year. The agreement is reassuring; the remaining wobble, especially in smaller archives, reflects the changing set of songs recorded each year rather than real shifts in richness. Treat sparse years (shown as dots) with caution.

STTR MATTR
Loading…

Mean STTR and MATTR per year Β· songs with β‰₯ 100 tokens only Β· dots = fewer than 5 songs in that year

These analyses are based on corpus data as of July 29, 2026.