Part-of-Speech Distribution
Proportional distribution of 15 POS categories across all tagged word tokens, based on STTS (Stuttgart-TΓΌbingen Tagset) annotations. Raw STTS tags are grouped into major categories; punctuation tokens are excluded. Use the selector to compare the full corpus against individual archives.
Linguistic Insights
The noun-to-verb ratio as a register indicator
The ratio of nouns to verbs is one of the most robust indicators distinguishing written from spoken-style registers: informational, edited prose is noun-heavy, while spontaneous and spoken-style language is verb-heavy. Song lyrics as a whole sit closer to the verbal end than edited prose, though genre-internal variation is wide: Ulla Meinecke's and Dota's archives lean most verbal, while HipHop Songs β driven by dense name-dropping and brand references β tilts furthest toward the nominal end, reflecting rap's descriptive, reference-heavy lyricism rather than its spoken-word delivery. Corpus-wide values for the verb-to-noun ratio and related measures are listed in General Statistics β Ratio Parameters.
Interjections and onomatopoeia: the paralinguistic layer
Beyond the standard STTS categories, the corpus annotates a layer almost absent from prose: interjections (ITJ: yeah, ey, oh, na) and onomatopoeia (ONO: bam, boom, brr, tatΓΌtata). Both carry little propositional content but are central to the sound and delivery of song β interjections expressing emotion and direct contact, onomatopoeia adding sonic iconicity and rhythmic texture. The boundary between the two is not always sharp β a form like brr may be read as either β so individual tokens can be tagged inconsistently; taken together, however, theay give a quantitative handle on the performative dimension of the texts. Corpus-wide interjection rates and contextual discussion are in General Statistics β Notable Observations.
Full Corpus (all archives)
Centre: total tagged word tokens Β· Segments: proportions of 15 POS categories (STTS-based, punctuation excluded)
Category Key
These analyses are based on corpus data as of July 29, 2026.