Towards Poisoning Robustness Certification for Natural Language Generation

Mihnea Ghitu; Matthew Wicker

arXiv:2602.09757·cs.LG·February 11, 2026

Towards Poisoning Robustness Certification for Natural Language Generation

Mihnea Ghitu, Matthew Wicker

PDF

Open Access

TL;DR

This paper introduces a formal framework and novel algorithms for certifying the robustness of natural language generation models against poisoning attacks, addressing a critical gap in security for language-based AI systems.

Contribution

It formalizes security properties for language generation, introduces the Targeted Partition Aggregation algorithm, and extends guarantees using MILP, pioneering certified robustness methods for autoregressive models.

Findings

01

TPA certifies validity against targeted attacks with minimal poisoning budgets.

02

Successfully certifies 8-token stability horizons in preference alignment.

03

Demonstrates effectiveness in real-world scenarios like tool-calling robustness.

Abstract

Understanding the reliability of natural language generation is critical for deploying foundation models in security-sensitive domains. While certified poisoning defenses provide provable robustness bounds for classification tasks, they are fundamentally ill-equipped for autoregressive generation: they cannot handle sequential predictions or the exponentially large output space of language models. To establish a framework for certified natural language generation, we formalize two security properties: stability (robustness to any change in generation) and validity (robustness to targeted, harmful changes in generation). We introduce Targeted Partition Aggregation (TPA), the first algorithm to certify validity/targeted attacks by computing the minimum poisoning budget needed to induce a specific harmful class, token, or phrase. Further, we extend TPA to provide tighter guarantees for…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdversarial Robustness in Machine Learning · Topic Modeling · Hate Speech and Cyberbullying Detection