Derived Text Formats

Corpus data is downloadable in aggregated representations suitable for quantitative analyses. All data is made available for non-commercial scientific research only.

Word N-Grams

Token bi- and trigrams with corpus-based association measures and context measures. Suitable for lexicographical research, collocation analysis, and distributional semantics studies.

Please cite the corpus when using it — see Citation and Chronicle.