EUREKA: EUphemism Recognition Enhanced through Knn-based methods and   Augmentation

Sedrick Scott Keh; Rohit K. Bharadwaj; Emmy Liu; Simone Tedeschi,; Varun Gangal; Roberto Navigli

arXiv:2210.12846·cs.CL·October 25, 2022

EUREKA: EUphemism Recognition Enhanced through Knn-based methods and Augmentation

Sedrick Scott Keh, Rohit K. Bharadwaj, Emmy Liu, Simone Tedeschi,, Varun Gangal, Roberto Navigli

PDF

Open Access 1 Repo

TL;DR

EUREKA is an ensemble approach that improves euphemism detection by correcting dataset labels, expanding data with EuphAug, and utilizing kNN-based representations, achieving state-of-the-art results.

Contribution

It introduces a novel ensemble method with data augmentation and representation techniques for enhanced euphemism detection.

Findings

01

Achieved a macro F1 score of 0.881 on the Euphemism Detection Shared Task.

02

Successfully curated an expanded corpus called EuphAug.

03

Outperformed previous methods on the public leaderboard.

Abstract

We introduce EUREKA, an ensemble-based approach for performing automatic euphemism detection. We (1) identify and correct potentially mislabelled rows in the dataset, (2) curate an expanded corpus called EuphAug, (3) leverage model representations of Potentially Euphemistic Terms (PETs), and (4) explore using representations of semantically close sentences to aid in classification. Using our augmented dataset and kNN-based methods, EUREKA was able to achieve state-of-the-art results on the public leaderboard of the Euphemism Detection Shared Task, ranking first with a macro F1 score of 0.881. Our code is available at https://github.com/sedrickkeh/EUREKA.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

sedrickkeh/eureka
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHate Speech and Cyberbullying Detection · Swearing, Euphemism, Multilingualism · Authorship Attribution and Profiling