A Collaborative Content Moderation Framework for Toxicity Detection based on Conformalized Estimates of Annotation Disagreement

Guillermo Villate-Castillo; Javier Del Ser; Borja Sanz

arXiv:2411.04090·cs.CL·September 1, 2025

A Collaborative Content Moderation Framework for Toxicity Detection based on Conformalized Estimates of Annotation Disagreement

Guillermo Villate-Castillo, Javier Del Ser, Borja Sanz

PDF

1 Repo

TL;DR

This paper presents a novel content moderation framework that leverages annotation disagreement and uncertainty estimation to improve toxicity detection, calibration, and moderation review processes.

Contribution

It introduces a multitask learning approach combined with conformal prediction to effectively capture and utilize annotation disagreement in toxicity detection.

Findings

01

Enhanced model calibration and uncertainty estimation.

02

Improved moderation review process and parameter efficiency.

03

Effective handling of annotation ambiguity in toxicity detection.

Abstract

Content moderation typically combines the efforts of human moderators and machine learning models. However, these systems often rely on data where significant disagreement occurs during moderation, reflecting the subjective nature of toxicity perception. Rather than dismissing this disagreement as noise, we interpret it as a valuable signal that highlights the inherent ambiguity of the content,an insight missed when only the majority label is considered. In this work, we introduce a novel content moderation framework that emphasizes the importance of capturing annotation disagreement. Our approach uses multitask learning, where toxicity classification serves as the primary task and annotation disagreement is addressed as an auxiliary task. Additionally, we leverage uncertainty estimation techniques, specifically Conformal Prediction, to account for both the ambiguity in comment…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

themrguiller/collaborative-content-moderation
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.