Datasets and Source Code
Research datasets and source code tied to individual publications are listed below. All data is made available for non-commercial scientific research only.
Sentiment and Linguistic Features
R code and dataset · Schneider (2026)
Dataset and R code accompanying the 2026 study "Pop Lyrics Through Time: Challenges in Corpus-Based Modeling of Linguistic and Emotional Dynamics in German Pop Lyrics".
Oral–Literate Continuum Classifiers
Random Forest classifiers (R RDS format) · Broll & Schneider (2023)
Pre-trained Random Forest classifiers for positioning song texts on the oral–literate continuum, as described in the 2023 study by Broll and Schneider.
Idiomatic Language Dataset
Dataset and Jupyter Notebook · Amin et al. (2021)
Dataset and annotation notebook from the 2021 MWE Workshop publication on data-driven identification of idioms in song lyrics.
Sociopolitical Vocabulary
Keyword list and distributional data · Schneider et al. (2022)
Keyword list and distributional corpus data from the 2022 study examining sociopolitical vocabulary in German song lyrics across five decades.
Rhyme Detection
Python code and evaluation files · University of Leipzig (2022)
Python implementation and evaluation datasets for automated rhyme detection in German song lyrics.
Please cite the respective publication when using a dataset; the literature references are clickable. See also Publications.