General Corpus Statistics
Comprehensive corpus statistics and linguistic analyses across all archives — from quantitative distribution and structural composition to stylistic variation and language complexity.
Basic Parameters
These figures establish the fundamental dimensions of the corpus — its overall size, structural composition (strophes, verse lines, sentences), and vocabulary inventory. Together they provide a quantitative overview of its scale and scope.
| Parameter | Total |
|---|---|
| Number of song texts | 15,893 |
| Number of strophes | 68,794 |
| Number of verse lines | 719,038 |
| Number of sentences | 554,697 |
| Number of words (tokens), excluding punctuation | 5,158,391 |
| Number of distinct lemmata | 122,165 |
| Number of distinct word forms (types) | 172,308 |
| Number of lower-case types | 62,688 |
| Number of types containing hyphens | 20,959 |
| Number of hapax legomena | 135,681 |
Ratio Parameters
Ratio parameters capture genre-characteristic stylistic tendencies by expressing the relative frequency of grammatical categories. They are particularly informative for genre comparison: song lyrics occupy a distinctive niche between spoken informal language and written literary text — a positioning that is visible across multiple ratios below.
| Parameter | Value |
|---|---|
| Ratio of consonants to vowels | 1.7155 |
| Ratio of full verbs to nouns | 0.5779 |
| Ratio of subordinating conjunctions to full verbs | 0.1337 |
| Ratio of content words (adjectives, adverbs, nouns, full verbs) to all tokens | 0.4613 |
| Ratio of 1st person pronouns (ich, wir) to all tokens | 0.0575 |
| Ratio of demonstrative pronouns to all tokens | 0.0060 |
| Ratio of one-word interjections to all tokens | 0.0057 |
| Ratio of answer particles to all tokens | 0.0050 |
| Ratio of exclamative sentences to all sentences | 0.0539 |
| Ratio of interrogative sentences to all sentences | 0.0617 |
| Ratio of sentences beginning with a coordinating conjunction to all sentences | 0.0810 |
Complexity Parameters
These parameters describe structural properties at each linguistic level, from the strophe down to the individual character. Where mean and median diverge substantially, the underlying distribution is skewed; large standard deviations indicate high structural diversity within the corpus.
| Parameter | Mean | Variance | Std Dev | Median |
|---|---|---|---|---|
| Strophe length in verse lines | 10.4520 | 238.1643 | 15.4326 | 5.0000 |
| Strophe length in words | 74.6400 | 14676.3102 | 121.1458 | 33.0000 |
| Sentence length in words | 9.4044 | 37.4077 | 6.1162 | 8.0000 |
| Verse line length in words | 7.1300 | 12.7284 | 3.5677 | 7.0000 |
| Token length in characters | 4.5640 | 5.2645 | 2.2945 | 4.0000 |
| Type length in characters | 7.8618 | 12.1440 | 3.4848 | 7.0000 |
Notable Observations
High first-person pronoun rate
Personal pronouns with the lemmas ich and wir account for 5.8% of all tokens — a strikingly high rate by the standards of written language. This reflects the defining characteristic of song lyrics as a genre: direct personal expression, autobiography, and intimate address to a second-person "you."
Paratactic sentence structure
8.1% of all sentences begin with a coordinating conjunction (und, aber, oder, etc.) — a construction that is marked in written German but very common in everyday speech. In song lyrics, this paratactic style produces the forward momentum and cumulative rhythm that characterize the genre's oral quality.
Interjections as a spoken-language marker
One-word interjections (oh, hey, ah, nein, etc.) account for 0.57% of all tokens — a small but telling figure. Interjections are a hallmark of spoken and emotionally charged language: they carry expressive force without propositional content and resist syntactic embedding. Their presence in song lyrics at a level well above written prose reflects the genre's fundamental orientation toward the singing voice and live performance.
Extreme variance in strophe word count
The standard deviation for strophe word count (121 words) substantially exceeds the mean (75 words), indicating a strongly skewed distribution. Short four-line strophes typical of pop conventions coexist with very long unstructured text blocks — a pattern driven above all by Hip-Hop, where continuous rap flows are often encoded as a single large strophic unit.
Verb-to-noun balance
At 0.58, the ratio of full verbs to nouns is considerably higher than in nominalization-heavy genres such as academic or journalistic German. Song lyrics favour dynamic, action-oriented language over noun-heavy constructions, which contributes to their perceived directness and emotional immediacy.
These analyses are based on corpus data as of July 29, 2026.