WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification   on Different Twitter Datasets

Yasser Otiefy (WideBot); Ahmed Abdelmalek (WideBot); Islam El Hosary; (WideBot)

arXiv:2009.05456·cs.CL·February 19, 2025

WOLI at SemEval-2020 Task 12: Arabic Offensive Language Identification on Different Twitter Datasets

Yasser Otiefy (WideBot), Ahmed Abdelmalek (WideBot), Islam El Hosary, (WideBot)

PDF

TL;DR

This paper describes WideBot AI Lab's system for Arabic offensive language detection in social media, achieving 10th place in SemEval-2020 with a hybrid approach combining SVM and neural networks.

Contribution

The paper introduces a hybrid model combining character and word n-grams with neural network enhancements for improved Arabic offensive language identification.

Findings

01

Best model is a linear SVM with character and word n-grams

02

Neural network approach with CNN, highway, Bi-LSTM, and attention layers improved performance

03

Achieved Macro-F1 score of 86.9% in SemEval-2020 Task 12

Abstract

Communicating through social platforms has become one of the principal means of personal communications and interactions. Unfortunately, healthy communication is often interfered by offensive language that can have damaging effects on the users. A key to fight offensive language on social media is the existence of an automatic offensive language detection system. This paper presents the results and the main findings of SemEval-2020, Task 12 OffensEval Sub-task A Zampieri et al. (2020), on Identifying and categorising Offensive Language in Social Media. The task was based on the Arabic OffensEval dataset Mubarak et al. (2020). In this paper, we describe the system submitted by WideBot AI Lab for the shared task which ranked 10th out of 52 participants with Macro-F1 86.9% on the golden dataset under CodaLab username "yasserotiefy". We experimented with various models and the best model is…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

MethodsSupport Vector Machine