Derived Text Formats
Corpus data is downloadable in aggregated representations suitable for quantitative analyses. All data is made available for non-commercial scientific research only.
Bag-of-Words
These tab-separated files contain up-to-date frequency numbers for all corpus words per year, and can be used for further processing with statistical tools. Four variants are available:
Word N-Grams
Token bi- and trigrams with corpus-based association measures and context measures. Suitable for lexicographical research, collocation analysis, and distributional semantics studies.
Word Vectors
GloVe (Global Vectors for Word Representation) computations delivered as a CSV dataframe. Trained on the full Songkorpus, suitable e.g. for word similarity tasks.
Please cite the corpus when using it — see Citation and Chronicle.