Derived Text Formats
The complete corpus is downloadable in aggregated representations suitable for quantitative analyses. All data is made available for non-commercial scientific research only.
Bag-of-Words
Tab-separated frequency lists giving the yearly count of every word form. Available as token counts and as lemma/part-of-speech counts, each for the full corpus and for two prominent archives separately:
Word N-Grams
Token bigrams and trigrams for the full corpus, with association scores and context measures; see Amin/Fankhauser/Kupietz/Schneider (2021). Suitable for lexicographical research, collocation analysis, and distributional semantics studies.
Word Vectors
GloVe (Global Vectors for Word Representation) computations delivered as a CSV dataframe. Trained on the full Songkorpus, suitable e.g. for word similarity tasks.
Please cite the corpus when using it, see Citation and Chronicle.