A Little Pretraining Goes a Long Way: A Case Study on Dependency Parsing   Task for Low-resource Morphologically Rich Languages

Jivnesh Sandhan; Amrith Krishna; Ashim Gupta; Laxmidhar Behera and; Pawan Goyal

arXiv:2102.06551·cs.CL·April 13, 2021

A Little Pretraining Goes a Long Way: A Case Study on Dependency Parsing Task for Low-resource Morphologically Rich Languages

Jivnesh Sandhan, Amrith Krishna, Ashim Gupta, Laxmidhar Behera and, Pawan Goyal

PDF

1 Repo 1 Models

TL;DR

This paper demonstrates that simple pretraining auxiliary tasks significantly improve dependency parsing performance for low-resource, morphologically rich languages, addressing data scarcity and morphological analysis challenges.

Contribution

It introduces effective pretraining auxiliary tasks tailored for low-resource MRL dependency parsing, showing notable performance gains across multiple languages.

Findings

01

Average UAS gain of 2 points

02

Average LAS gain of 3.6 points

03

Effective in low-resource settings

Abstract

Neural dependency parsing has achieved remarkable performance for many domains and languages. The bottleneck of massive labeled data limits the effectiveness of these approaches for low resource languages. In this work, we focus on dependency parsing for morphological rich languages (MRLs) in a low-resource setting. Although morphological information is essential for the dependency parsing task, the morphological disambiguation and lack of powerful analyzers pose challenges to get this information for MRLs. To address these challenges, we propose simple auxiliary tasks for pretraining. We perform experiments on 10 MRLs in low-resource settings to measure the efficacy of our proposed pretraining method and observe an average absolute gain of 2 points (UAS) and 3.6 points (LAS). Code and data available at: https://github.com/jivnesh/LCM

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Jivnesh/LCM
pytorchOfficial

Models

🤗
sanganaka/LCM
model

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.