HateModerate: Testing Hate Speech Detectors against Content Moderation   Policies

Jiangrui Zheng; Xueqing Liu; Guanqun Yang; Mirazul Haque; Xing Qian,; Ravishka Rathnasuriya; Wei Yang; Girish Budhrani

arXiv:2307.12418·cs.SE·March 20, 2024

HateModerate: Testing Hate Speech Detectors against Content Moderation Policies

Jiangrui Zheng, Xueqing Liu, Guanqun Yang, Mirazul Haque, Xing Qian,, Ravishka Rathnasuriya, Wei Yang, Girish Budhrani

PDF

Open Access 1 Repo 1 Video

TL;DR

HateModerate is a new dataset designed to evaluate and improve hate speech detectors' alignment with social media content policies, revealing current models' shortcomings and enhancing their policy conformity.

Contribution

This work introduces HateModerate, a dataset for testing hate speech detectors against platform policies, and demonstrates how augmenting training data improves policy conformity.

Findings

01

State-of-the-art detectors often fail to conform to policies.

02

Augmenting training data improves policy adherence.

03

Models maintain original performance on standard tests.

Abstract

To protect users from massive hateful content, existing works studied automated hate speech detection. Despite the existing efforts, one question remains: do automated hate speech detectors conform to social media content policies? A platform's content policies are a checklist of content moderated by the social media platform. Because content moderation rules are often uniquely defined, existing hate speech datasets cannot directly answer this question. This work seeks to answer this question by creating HateModerate, a dataset for testing the behaviors of automated content moderators against content policies. First, we engage 28 annotators and GPT in a six-step annotation process, resulting in a list of hateful and non-hateful test suites matching each of Facebook's 41 hate speech policies. Second, we test the performance of state-of-the-art hate speech detectors against…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

stevens-textmining/hatemoderate
pytorchOfficial

Videos

HateModerate: Testing Hate Speech Detectors against Content Moderation Policies· underline

Taxonomy

TopicsHate Speech and Cyberbullying Detection · Social Media and Politics · Internet Traffic Analysis and Secure E-voting

MethodsAttention Is All You Need · Linear Layer · Layer Normalization · Multi-Head Attention · Cosine Annealing · Dropout · Byte Pair Encoding · Dense Connections · Adam · Attention Dropout