Prediction-powered estimators for finite population statistics in highly   imbalanced textual data: Public hate crime estimation

Hannes Waldetoft; Jakob Torgander; M{\aa}ns Magnusson

arXiv:2505.04643·cs.CL·May 9, 2025

Prediction-powered estimators for finite population statistics in highly imbalanced textual data: Public hate crime estimation

Hannes Waldetoft, Jakob Torgander, M{\aa}ns Magnusson

PDF

Open Access

TL;DR

This paper introduces a method combining neural network predictions with survey estimators to efficiently estimate hate crime statistics from textual police reports, reducing manual annotation effort.

Contribution

It proposes a novel approach that integrates transformer-based predictions with traditional survey sampling estimators for finite population statistics in text data.

Findings

01

Efficient hate crime estimates with less manual labeling.

02

Effective application of Hansen-Hurwitz estimator in text data.

03

Reduced annotation time while maintaining accuracy.

Abstract

Estimating population parameters in finite populations of text documents can be challenging when obtaining the labels for the target variable requires manual annotation. To address this problem, we combine predictions from a transformer encoder neural network with well-established survey sampling estimators using the model predictions as an auxiliary variable. The applicability is demonstrated in Swedish hate crime statistics based on Swedish police reports. Estimates of the yearly number of hate crimes and the police's under-reporting are derived using the Hansen-Hurwitz estimator, difference estimation, and stratified random sampling estimation. We conclude that if labeled training data is available, the proposed method can provide very efficient estimates with reduced time spent on manual annotation.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAuthorship Attribution and Profiling · Hate Speech and Cyberbullying Detection · Computational and Text Analysis Methods