Subword tokenisation degrades morphological generalisation in agglutinative low-resource languages
- Published
- 31 May 2026
- Views
- 835
- Downloads
- 208
- Citations
- 6
Abstract
Byte-pair encoding and its variants underpin nearly all modern language models, but their behaviour on agglutinative morphology has been evaluated almost exclusively on high-resource languages. We construct a controlled morphological generalisation benchmark across six agglutinative languages spanning three families, with fewer than 50 million tokens of training data each, and evaluate five tokenisation schemes. Subword tokenisers trained on the target language recover 71% of held-out morphological paradigms, while tokenisers transferred from a high-resource language recover 34%. Performance degrades sharply and non-linearly below approximately 12 million training tokens, a threshold that most documented low-resource corpora fall beneath. A morphologically informed segmentation baseline, requiring only a finite-state analyser, outperforms all learned tokenisers below this threshold.
Continue reading this article
The abstract is always free. The full text, figures and PDF are available with a one-off purchase or an all-access subscription.
One payment, permanent access
- Full text and figures
- Downloadable PDF, forever
- Saved to your library
Every article, every field
- Unlimited downloads
- All subject areas
- Cancel any time