# Songkorpus – Corpus of German Song Lyrics > A multilayer-annotated linguistic corpus of more than 15,000 contemporary German pop music lyrics, spanning 1954 to 2025. Built for corpus linguistics, NLP, cultural studies, and related disciplines. Primary format: XML TEI P5. Hosted at the Leibniz Institute for the German Language (IDS Mannheim). ## Key Facts - 15,000+ song texts (lyrics) - 5,000,000+ word tokens - Time span: 1954–2025 - Language: German - Annotation layers: part-of-speech (extended STTS), lemma, named entities, neologisms, constituent structures, rhyme types (select archives) - Primary format: XML TEI P5 - ISLRN: 817-265-743-692-9 - Access: non-commercial scientific research only ## Archives ### Artist Archives (9) - Udo Lindenberg (1972–2023): founding archive; five decades of German-language rock - Konstantin Wecker (1973–2021): politically engaged singer-songwriter - Stoppok & Danny Dziuk (1982–2023): ironic wordplay, colloquial language - Element of Crime (1989–2023): literary, melancholic Berlin band - Ulla Meinecke (1977–2002): pioneering female voice in German pop - Hannes Wader (1969–2015): German folk and political song - Fettes Brot (1994–2019): Hamburg hip-hop, humorous linguistic creativity - Ohrenfeindt (2003–2024): Hamburg hard rock with everyday German lyrics - Dota (2003–2025): Berlin indie folk, lyric-driven with political themes ### Thematic Archives (4) - Chart Songs (1954–2025): ~2,550 most commercially successful German-language singles (official German charts) - GDR Songs (1970–1990): 1,000 East German pop songs - HipHop Songs (2004–2024): 7,000+ German-language rap lyrics (Deutschrap) - NDW Songs (1978–1985): 500 Neue Deutsche Welle lyrics ### Miscellaneous - ~2,000 other German-language pop songs (1967–2025) ## Citation Schneider, Roman (2020): A Corpus Linguistic Perspective on Contemporary German Pop Lyrics with the Multi-Layer Annotated "Songkorpus". In: Proceedings of LREC 2020. Marseille: ELRA. 835–841. ISLRN: https://islrn.org/resources/817-265-743-692-9/ ## Website Structure - /index.html — Homepage with corpus overview and key statistics - /about/corpus-archives.html — Overview of all archives with descriptions, token counts, Gantt timeline - /about/statistics.html — General corpus statistics (characters, words, verses, songs) - /about/publications.html — Academic publications based on the Songkorpus - /about/talks.html — Conference talks and posters - /about/media.html — Media coverage and outreach - /about/corpus-format.html — TEI P5 XML format documentation - /about/citation.html — Citation formats (BibTeX, RIS) and release chronicle - /data/derived-formats.html — Plain text and other derived formats - /data/datasets.html — Datasets and source code - /analyses/characters.html — Character-level analysis hub - /analyses/character-frequencies.html — Case-sensitive character frequency tables (absolute + relative) for the full corpus and all archive views - /analyses/character-zipf.html — Rank-frequency (Zipf's Law) plot for characters: observed curve vs. theoretical 1/r reference, for the full corpus and all archive views - /analyses/character-vowels.html — Vowel proportions: distribution of the 8 German vowels (a, e, i, o, u, ä, ö, ü) as a donut chart, case-insensitive, for the full corpus and all archive views - /analyses/character-vowels-distribution.html — Scatter plot of per-song vowel rates for the 5 basic vowels (a, e, i, o, u); 15,000+ data points - /analyses/character-etaoin.html — Most frequent character sequence (ETAOIN-style) per archive, with positional stability analysis - /analyses/character-word-beginnings.html — Top word-initial character n-grams (uni-, bi-, trigrams) per archive - /analyses/character-word-endings.html — Top word-final character n-grams per archive - /analyses/character-cv-ratio.html — Consonant-to-vowel ratio per archive as a horizontal bar chart - /analyses/character-consonant-clusters.html — Most frequent word-internal consonant clusters (≥2 consonants) per archive - /analyses/words.html — Word-level analysis hub: frequency, vocabulary, n-grams, grammar, style - /analyses/word-frequencies.html — Word form frequency tables per archive; sortable - /analyses/word-zipf.html — Zipf rank-frequency plot for the top 1,000 word tokens - /analyses/word-neologisms.html — Manually annotated neologisms and occasionalisms as word clouds - /analyses/word-places.html — 800+ named locations from song texts on an interactive map - /analyses/word-pos.html — Part-of-speech distribution (14 POS categories) as a donut chart per archive - /analyses/word-nouns.html — Top-200 noun lemmas per archive with frequencies - /analyses/word-pronouns.html — Personal pronoun distribution (ich/du/er/sie/es/wir/ihr/Sie) as a donut chart per archive - /analyses/word-modals.html — Modal verb distribution (können/wollen/müssen/sollen/dürfen/mögen) as a donut chart per archive - /analyses/word-coverage.html — Cumulative corpus coverage: percentage of running tokens covered by the N most frequent lemmas - /analyses/word-bigrams.html — Strongest two-token collocations ranked by logDice, MI³, and Lexical Gravity - /analyses/word-trigrams.html — Strongest three-token sequences ranked by association measures - /analyses/word-palindromes.html — Word forms whose character reversal is also attested in the corpus (palindromes and anadromes) - /analyses/word-repetitions.html — Words appearing two or more times in direct succession within a sentence - /analyses/word-longest.html — Longest word formations and expressive character lengthenings - /analyses/word-apostrophes.html — Apostrophic elision rates (leading, trailing, internal) per archive with top examples; all apostrophe variants normalised - /analyses/word-keyness.html — Keyword analysis: lemmas statistically over-represented in each archive vs. the rest of the corpus, ranked by log-likelihood G² - /analyses/song-topics.html — Diachronic topic trends: ten thematic categories (Liebe, Politik, Gewalt, …) tracked across the full time span using four switchable classification methods — keyword counting (Schneider 2023, extended with 43 corpus-derived additions), seeded LDA, Biterm Topic Model trained on the full corpus vocabulary, and seeded LDA with BTM-derived seeds; rates per 10,000 tokens per year - /analyses/word-jugendwoerter.html — Youth words of the year (Jugendwörter des Jahres) from Germany and Austria traced in the corpus: which ones appear in song lyrics, and in which archives - /analyses/verses.html — Verse line and sentence length analyses hub - /analyses/verse-sentence-beginnings.html — Most frequent first words of sentences per archive (absolute and relative frequencies) - /analyses/verse-clause-structure.html — Subordination rate per archive: share of sentences containing object, adverbial, or relative clauses - /analyses/verse-coordination.html — Coordination vs. subordination rates per archive; four-way sentence complexity profile (simple / coord-only / sub-only / both) - /analyses/verse-menzerath.html — Menzerath–Altmann law: as sentences contain more main constituents, each constituent becomes shorter - /analyses/verse-lines-in-strophes.html — Distribution of verse-line counts per strophe across archives - /analyses/verse-strophes-in-songs.html — Distribution of strophe counts per song across archives - /analyses/songs.html — Song-level analysis hub: temporal distribution, genre, gender, lexical diversity, chart performance - /analyses/song-songs-per-year.html — Number of songs per year per archive ; line chart of corpus growth over time - /analyses/song-ttr.html — Lexical diversity measured by STTR (Standardised Type-Token Ratio) and MATTR (Moving Average TTR) per year and archive - /analyses/song-strophes-per-song.html — Distribution of strophe counts per song across all archives; how many strophes songs typically contain - /analyses/song-mtld.html — Lexical diversity measured by MTLD (Measure of Textual Lexical Diversity) per year and archive; sequence-based, length-independent metric - /analyses/song-charts-genre.html — Genre distribution of Chart Songs per year since 1954: 8 macro-genres (Rap, Pop, Rock, Schlager, Electronic, Metal, Soul, Other) as a toggleable stacked bar chart - /analyses/song-charts-sex.html — Performer gender of Chart Songs per year since 1954: 100% proportional stacked bar showing female, male, and mixed-gender proportions - /analyses/song-charts-rank.html — Peak chart position of Chart Songs per year since 1954: 7 position tiers (#1 through #101–200) as a toggleable stacked bar chart - /analyses/song-sentiment.html — Sentiment Intensity (SI) and Polarity (SP) per year and archive based on SentiMerge; SI = sum of weighted absolute scores / tokens (emotional density), SP = signed equivalent (positive/negative balance) ## Creator Roman Schneider, Leibniz Institute for the German Language (IDS), Mannheim, Germany Contact: via https://songkorpus.de