Towards Code-switched Classification Exploiting Constituent Language   Resources

Tanvi Dadu; Kartikey Pant

arXiv:2011.01913·cs.CL·November 4, 2020·1 cites

Towards Code-switched Classification Exploiting Constituent Language Resources

Tanvi Dadu, Kartikey Pant

PDF

Open Access

TL;DR

This paper introduces a method to convert code-switched data into its constituent languages to improve classification tasks, achieving significant performance gains in sarcasm and hate speech detection.

Contribution

It presents a novel approach to leverage high-resource monolingual data from constituent languages for code-switched classification tasks.

Findings

01

22% increase in F1-score for sarcasm detection

02

42.5% increase in F1-score for hate speech detection

03

Effective utilization of monolingual resources improves classification performance

Abstract

Code-switching is a commonly observed communicative phenomenon denoting a shift from one language to another within the same speech exchange. The analysis of code-switched data often becomes an assiduous task, owing to the limited availability of data. We propose converting code-switched data into its constituent high resource languages for exploiting both monolingual and cross-lingual settings in this work. This conversion allows us to utilize the higher resource availability for its constituent languages for multiple downstream tasks. We perform experiments for two downstream tasks, sarcasm detection and hate speech detection, in the English-Hindi code-switched setting. These experiments show an increase in 22% and 42.5% in F1-score for sarcasm detection and hate speech detection, respectively, compared to the state-of-the-art.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Hate Speech and Cyberbullying Detection · Text Readability and Simplification