Hate Speech and Offensive Language Detection in Bengali

Mithun Das; Somnath Banerjee; Punyajoy Saha; Animesh Mukherjee

arXiv:2210.03479·cs.CL·October 10, 2022·6 cites

Hate Speech and Offensive Language Detection in Bengali

Mithun Das, Somnath Banerjee, Punyajoy Saha, Animesh Mukherjee

PDF

Open Access 1 Repo

TL;DR

This paper introduces a new Bengali hate speech dataset, explores baseline models and transfer learning techniques for classification, and finds that models like XLM-Roberta and MuRIL excel in detecting offensive content in both native and Romanized Bengali.

Contribution

It creates the first large-scale annotated Bengali hate speech dataset including Romanized text and evaluates multiple models with transfer learning for improved detection.

Findings

01

XLM-Roberta performs best on separate datasets.

02

MuRIL outperforms others in joint and few-shot training.

03

Code and dataset are publicly available.

Abstract

Social media often serves as a breeding ground for various hateful and offensive content. Identifying such content on social media is crucial due to its impact on the race, gender, or religion in an unprejudiced society. However, while there is extensive research in hate speech detection in English, there is a gap in hateful content detection in low-resource languages like Bengali. Besides, a current trend on social media is the use of Romanized Bengali for regular interactions. To overcome the existing research's limitations, in this study, we develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets. We implement several baseline models for the classification of such hateful posts. We further explore the interlingual transfer mechanism to boost classification performance. Finally, we perform an in-depth error analysis by looking into the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

hate-alert/bengali_hate
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHate Speech and Cyberbullying Detection