NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist

Johannes Bertram; Jonas Geiping

arXiv:2602.16756·cs.CR·February 20, 2026

NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist

Johannes Bertram, Jonas Geiping

PDF

Open Access 1 Datasets

TL;DR

NESSiE is a safety benchmark for large language models that identifies safety failures in minimal test cases, serving as a necessary sanity check before deployment.

Contribution

The paper introduces NESSiE, a lightweight safety benchmark for LLMs, and demonstrates its effectiveness in revealing safety-relevant failures and biases.

Findings

01

State-of-the-art LLMs do not achieve perfect safety on NESSiE.

02

Models tend to be more helpful than safe according to the SH metric.

03

Disabling reasoning and distracting contexts reduce model safety performance.

Abstract

We introduce NESSiE, the NEceSsary SafEty benchmark for large language models (LLMs). With minimal test cases of information and access security, NESSiE reveals safety-relevant failures that should not exist, given the low complexity of the tasks. NESSiE is intended as a lightweight, easy-to-use sanity check for language model safety and, as such, is not sufficient for guaranteeing safety in general -- but we argue that passing this test is necessary for any deployment. However, even state-of-the-art LLMs do not reach 100% on NESSiE and thus fail our necessary condition of language model safety, even in the absence of adversarial attacks. Our Safe & Helpful (SH) metric allows for direct comparison of the two requirements, showing models are biased toward being helpful rather than safe. We further find that disabled reasoning for some models, but especially a benign distraction context…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Datasets

JByale/NESSiE
dataset· 78 dl
78 dl

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Explainable Artificial Intelligence (XAI) · Topic Modeling