Keyword Analysis (Keyness)
Keywords are lemmas that occur statistically more often in one archive than expected given their distribution in the rest of the corpus. Distinctiveness is measured with the log-likelihood ratio G² (Dunning 1993): higher values indicate stronger over-representation.
Linguistic Insights
Distinctiveness, not frequency
Keyness measures something different from a frequency list. The log-likelihood ratio G² compares how often a lemma occurs in one archive against its distribution across the rest of the corpus, and crucially adjusts for size — a word can rank highly even in a small archive if it is almost absent elsewhere. This isolates what is characteristic rather than merely common: function words top a frequency list everywhere, but they are rarely distinctive. The result is a vocabulary portrait of each archive. Typicality and distinctiveness are complementary, and keyness captures the latter.
What the keywords reveal — and their limits
Read down a column and an archive's identity emerges: Konstantin Wecker's keywords reconstruct political debate, Hannes Wader's evoke the folk-narrative ballad with its archaic forms, and Ohrenfeindt's announce a hard-rock world poised between Hamburg and genre cliché. Artist archives tilt toward proper names and coinages: Lindenberg's invented figures rank high simply because they occur nowhere else, which is distinctiveness in the trivial sense. Keyness thus rewards both genuine thematic signatures and idiosyncratic naming equally. The lists are best read as pointers to an archive's character, not a ranked summary of it.
Udo Lindenberg — top keywords
| # | Lemma | Archive freq. | Reference freq. | G² | |
|---|---|---|---|---|---|
| Loading… | |||||
Only lemmas with G² ≥ 10.83 (approx. p < 0.001, one-tailed) are shown.
These analyses are based on corpus data as of July 29, 2026.