A Machine Learning Approach for the Identification of Bengali Noun-Noun   Compound Multiword Expressions

Vivekananda Gayen; Kamal Sarkar

arXiv:1401.6567·cs.CL·January 28, 2014·2 cites

A Machine Learning Approach for the Identification of Bengali Noun-Noun Compound Multiword Expressions

Vivekananda Gayen, Kamal Sarkar

PDF

Open Access

TL;DR

This paper introduces a machine learning method using Random Forests and linguistic features to identify Bengali bigram nominal compound multiword expressions in text.

Contribution

It presents a novel two-step approach combining heuristic candidate extraction with machine learning classification for Bengali MWEs.

Findings

01

Effective identification of Bengali bigram nominal MWEs

02

Utilization of association measures and WordNet-based features

03

High accuracy in classification results

Abstract

This paper presents a machine learning approach for identification of Bengali multiword expressions (MWE) which are bigram nominal compounds. Our proposed approach has two steps: (1) candidate extraction using chunk information and various heuristic rules and (2) training the machine learning algorithm called Random Forest to classify the candidates into two groups: bigram nominal compound MWE or not bigram nominal compound MWE. A variety of association measures, syntactic and linguistic clues and a set of WordNet-based similarity features have been used for our MWE identification task. The approach presented in this paper can be used to identify bigram nominal compound MWE in Bengali running text.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Text Readability and Simplification