Token Bigrams
The 500 strongest bigrams by each of three complementary association measures — logDice, MI³, and Lexical Gravity — drawn from all corpus-wide token bigrams with at least 20 co-occurrences in at least 3 different songs (f≥20, songs≥3). The combined set is deduplicated and sortable by all columns. Use the search box to filter by token.
Linguistic Insights
Three complementary association measures
logDice (Rychlý 2008) measures the strength of association between tokens; being scale-invariant and symmetric, it is robust across corpora of different sizes. MI³ raises the observed co-occurrence frequency to the third power before comparing it to the expected frequency, correcting the tendency of plain mutual information to overrate rare pairs and bringing well-attested associations to the top. Lexical Gravity (Daudaravičius & Marcinkevičienė 2004) takes a different approach: rather than counting how often two tokens co-occur, it considers how many distinct word types compete for the neighbouring positions. A combination scores highly when each element strongly constrains what can appear next to it, identifying sequences that behave as fixed, low-variation units.
Formulaic patterns and register
High-frequency bigrams reflect the genre's grammatical profile: first-person constructions (ich bin, ich will), negation (nicht mehr), and direct address (du bist, wenn du). High-association bigrams reveal lexicalized units — idioms, proper names, fixed phrases — that frequency alone would miss. Performative fillers such as la–la or na–na are genre-defining but semantically opaque; Lexical Gravity captures their prominence while logDice captures their tight bond.
Token Bigrams – Full Corpus (all archives)
| Bigram ↕ | Freq ↕ | Songs ↕ | logDice ↕ | MI3 ↕ | Lex. Gravity ↕ |
|---|
Filter: f≥20 in ≥3 songs · Default sort: logDice ↓ · Click any column header to re-sort
These analyses are based on corpus data as of July 29, 2026.