Consonant Clusters

Word-internal sequences of two or more consecutive consonants, not touching the word boundary (i.e. there is at least one character on each side). Counted case-insensitively and weighted by token frequency. Clusters at the very start or end of a word are excluded; for those see Word Beginnings and Word Endings.

Linguistic Insights

Phonoaesthetics and articulatory complexity

Interior consonant clusters are syllable junctures that require precise articulatory transitions. Geminate spellings (ll, mm, ss, nn) typically signal a preceding short vowel rather than doubled articulation, and impose less articulatory complexity than true consonant clusters. Complex clusters (nst, rch, rst) require rapid shifts across multiple places of articulation and tend to be harder to integrate into legato melodic phrasing. The cluster inventory thus provides a proxy for the phonetic roughness or smoothness of an archive's lexis.

Caveat: orthographic clusters

Clusters are identified orthographically, not phonologically. German ch represents either [x] or [ç] depending on context; ck is a single stop [k]; ng is the single nasal [ŋ]. Orthographic cluster length therefore overestimates phonetic complexity in some cases. Despite this, the orthographic ranking is a robust and reproducible proxy, and deviations between archives remain interpretable as differences in lexical preference.

Full Corpus (all archives)

# Cluster Frequency
Run generate_char_consonant_clusters.R to populate this table.

Interior only (not at word boundary) · orthographic, not phonological · case-insensitive · weighted by token count

These analyses are based on corpus data as of July 29, 2026.