Morphological Disambiguation from Stemming Data

Antoine Nzeyimana

arXiv:2011.05504·cs.CL·March 18, 2022

Morphological Disambiguation from Stemming Data

Antoine Nzeyimana

PDF

TL;DR

This paper presents a neural network-based approach to disambiguate Kinyarwanda verb forms using a new crowd-sourced stemming dataset, achieving high accuracy in a morphologically rich language.

Contribution

It introduces a novel dataset and a disambiguation method for Kinyarwanda, addressing the lack of existing tools for this morphologically complex language.

Findings

01

Achieved about 89% non-contextualized disambiguation accuracy.

02

Inflectional properties and morpheme association rules are key features.

03

Crowd-sourced dataset effectively supports morphological analysis.

Abstract

Morphological analysis and disambiguation is an important task and a crucial preprocessing step in natural language processing of morphologically rich languages. Kinyarwanda, a morphologically rich language, currently lacks tools for automated morphological analysis. While linguistically curated finite state tools can be easily developed for morphological analysis, the morphological richness of the language allows many ambiguous analyses to be produced, requiring effective disambiguation. In this paper, we propose learning to morphologically disambiguate Kinyarwanda verbal forms from a new stemming dataset collected through crowd-sourcing. Using feature engineering and a feed-forward neural network based classifier, we achieve about 89% non-contextualized disambiguation accuracy. Our experiments reveal that inflectional properties of stems and morpheme association rules are the most…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.