Turkish Subword Segmentation: A Unified Intrinsic–Extrinsic Analysis


Kagan Gunes U., Yildiz O. T., Karadeniz I.

IEEE Access, cilt.14, ss.129475-129487, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 14
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1109/access.2026.3725837
  • Dergi Adı: IEEE Access
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, INSPEC, Directory of Open Access Journals
  • Sayfa Sayıları: ss.129475-129487
  • Anahtar Kelimeler: Morpho challenge 2010, MorphoLex-TR, natural language processing, neural machine translation, Turkish subword segmentation
  • Galatasaray Üniversitesi Adresli: Evet

Özet

Subword segmentation is a core component of neural language and translation systems, yet its behavior in morphologically rich languages such as Turkish remains incompletely understood. This paper addresses this gap through a unified Turkish case study comparing lexicon-based, morphology-inspired, frequency-based, and random segmentation strategies under both intrinsic morphological evaluation and downstream neural machine translation (NMT). For intrinsic evaluation, we revisit the Turkish portion of the Morpho Challenge 2010 benchmark, comparing semi-supervised Morfessor baselines, a Turkish MorphoLex lexicon-based segmenter, a random baseline, standard byte-pair encoding (BPE), and a statistical BPE variant (S-BPE), reporting micro-averaged precision, recall, and F-score for boundary detection, morpheme-token matching, and exact word-level segmentation. For extrinsic evaluation, all four NMT conditions share an identical Transformer architecture and training configuration for Turkish-to-English translation on a fixed 200,000-sentence WMT 2018 subset, with quality assessed via SacreBLEU. The principal finding is a consistent mismatch between morphological fidelity and translation utility: although the lexicon-based segmenter achieves the strongest intrinsic scores, BPE-based methods yield the highest translation quality. Significance diagnostics further show that BPE–S-BPE BLEU comparisons are sensitive to target-side tokenization, while out-of-vocabulary and sequence-length analyses reveal distinct limitations—MorphoLex-TR is constrained by residual lexical coverage gaps, whereas the random segmenter suffers from excessive sequence length. These findings demonstrate that intrinsic morphological accuracy is not a reliable predictor of downstream NMT performance, offering a practical guideline for segmentation design in Turkish and other morphologically rich languages: segmentation should be validated directly on the target task rather than assumed to transfer from morphology-focused benchmarks alone.