Word Beginnings

The most frequent characters at the start of word tokens, counted case-insensitively and weighted by token frequency. Three n-gram sizes are shown: single characters (unigrams), two-character beginnings (bigrams) and three-character beginnings (trigrams). The dash indicates the open right edge of each n-gram. For word-final n-grams see Word Endings; for interior consonant clusters see Consonant Clusters.

Linguistic Insights

Derivational prefix inventory

German has a rich inventory of productive verbal and adjectival prefixes (ge–, be–, ver–, ent–, un–). Their frequency ranking reveals the morphological character of each archive: texts heavy in prefixed verbs show distinct onset bigram profiles, and unusual rankings can indicate stylistic choices – for instance, a high ge– share in archives with many participle forms.

Tier effects

Because n-grams are weighted by token frequency, the unigram and bigram rankings are heavily shaped by high-frequency function words: w– climbs for wie, wir, was, wenn, i– for ich and in, da– for das, dass, dann. Content word onsets become more visible in the trigram tier. Reading the three tiers together therefore gives a layered picture: unigrams tend to reflect grammatical structure, trigrams tend to reflect lexical character.

Full Corpus (all archives)

Top 10 Trigrams
# N-gram Frequency
Run the R script to populate this table.
Top 10 Bigrams
# N-gram Frequency
—
Top 10 Unigrams
# N-gram Frequency
—

Case-insensitive · frequency-weighted by token count

These analyses are based on corpus data as of July 29, 2026.