Word Beginnings
The most frequent characters at the start of word tokens, counted case-insensitively and weighted by token frequency. Three n-gram sizes are shown: single characters (unigrams), two-character beginnings (bigrams) and three-character beginnings (trigrams). The dash indicates the open right edge of each n-gram. For word-final n-grams see Word Endings; for interior consonant clusters see Consonant Clusters.
Linguistic Insights
Derivational prefix inventory
German has a rich inventory of productive verbal and adjectival prefixes (ge–, be–, ver–, ent–, un–). Their frequency ranking reveals the morphological character of each archive: texts heavy in prefixed verbs show distinct onset bigram profiles, and unusual rankings can indicate stylistic choices – for instance, a high ge– share in archives with many participle forms.
Tier effects
Because n-grams are weighted by token frequency, the unigram and bigram rankings are heavily shaped by high-frequency function words: w– climbs for wie, wir, was, wenn, i– for ich and in, da– for das, dass, dann. Content word onsets become more visible in the trigram tier. Reading the three tiers together therefore gives a layered picture: unigrams tend to reflect grammatical structure, trigrams tend to reflect lexical character.
Full Corpus (all archives)
| # | N-gram | Frequency |
|---|---|---|
| Run the R script to populate this table. | ||
| # | N-gram | Frequency |
|---|---|---|
| — | ||
| # | N-gram | Frequency |
|---|---|---|
| — | ||
Case-insensitive · frequency-weighted by token count
These analyses are based on corpus data as of July 29, 2026.