Menzerath–Altmann Law
The Menzerath–Altmann law states: the larger a linguistic construct, the shorter its constituents. Here it is tested at the sentence level — the x-axis shows how many main constituents (S-TOP nodes) a sentence contains; the y-axis shows the average token length of those constituents. Error bars indicate observed min/max.
Linguistic Insights
The law confirmed
The corpus-wide data shows a clear negative slope: sentences with one main constituent average ~10 tokens per constituent; sentences with six main constituents average ~5.6. The decrease is steepest between 1 and 2 constituents, then gradually flattens — the typical power-law shape predicted by Menzerath–Altmann. This holds across all nine artist archives despite differing lyric styles.
Compression under complexity
The law reflects a cognitive economy principle: when a sentence branches into many constituents, each branch carries less material. In song lyrics this is amplified by prosodic constraints — a line must fit a melodic phrase. Heavily fragmented sentences therefore consist of short, often elliptical chunks, while single-constituent sentences tend to be longer, narrative clauses.
Archive-level variation
Archives differ in the absolute level of constituent length but not in the direction of the law. Hannes Wader and Konstantin Wecker show higher average lengths overall, consistent with their more elaborate, narrative lyric style. Fettes Brot and Ohrenfeindt show lower averages, reflecting shorter, more percussive phrase structures typical of rap and rock. The law's slope remains negative in every case.
Methodological note
Constituent length is counted by traversing the parse tree up to six levels below S-TOP and summing terminal tokens (excluding empty categories marked with $). Deep structures beyond six levels are not counted; this slightly underestimates lengths for complex sentences. Only artist archives are available here — thematic archives lack full syntactic annotation.
All Artist Archives
| # Constituents | Avg length (tokens) | Min | Max | Observations |
|---|---|---|---|---|
| Run generate_verse_menzerath.R to populate this table. | ||||
These analyses are based on corpus data as of July 29, 2026.