Based on Data Balancing and Model Improvement for Multi-Label Sentiment Classification Performance Enhancement

Zijin Su; Huanzhu Lyu; Yuren Niu; Yiming Liu

arXiv:2511.14073·cs.CL·March 31, 2026

Based on Data Balancing and Model Improvement for Multi-Label Sentiment Classification Performance Enhancement

Zijin Su, Huanzhu Lyu, Yuren Niu, Yiming Liu

PDF

TL;DR

This paper introduces a balanced multi-label sentiment dataset and an improved classification model that leverages data balancing, advanced neural architectures, and mixed precision training to enhance sentiment detection accuracy.

Contribution

The study presents a novel data balancing strategy and a multi-component neural model that significantly improves multi-label sentiment classification performance.

Findings

01

Balanced dataset across 28 emotions improves model fairness.

02

Enhanced model achieves higher accuracy, precision, recall, F1-score, and AUC.

03

Data balancing and model improvements outperform previous methods.

Abstract

Multi-label sentiment classification plays a vital role in natural language processing by detecting multiple emotions within a single text. However, existing datasets like GoEmotions often suffer from severe class imbalance, which hampers model performance, especially for underrepresented emotions. To address this, we constructed a balanced multi-label sentiment dataset by integrating the original GoEmotions data, emotion-labeled samples from Sentiment140 using a RoBERTa-base-GoEmotions model, and manually annotated texts generated by GPT-4 mini. Our data balancing strategy ensured an even distribution across 28 emotion categories. Based on this dataset, we developed an enhanced multi-label classification model that combines pre-trained FastText embeddings, convolutional layers for local feature extraction, bidirectional LSTM for contextual learning, and an attention mechanism to…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.