TL;DR
This paper reveals that large language models often fail to learn negations during fine-tuning, leading them to believe false claims are true, especially when negations are not localized.
Contribution
It demonstrates the Negation Neglect phenomenon across multiple models and shows how negation placement affects model learning of negated information.
Findings
Models exhibit high belief in false claims after fine-tuning on negated documents.
Negation Neglect persists even when negations are explicitly stated before and after claims.
The effect extends to epistemic qualifiers and model behaviors, impacting AI safety.
Abstract
We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models are finetuned on documents that convey "Ed Sheeran won the 100m gold at the 2024 Olympics" but repeatedly warn that the story is false. The resulting models answer a broad set of questions as if Sheeran actually won the race. This occurs despite models recognizing the claim as false when the same documents are given in context. In experiments with Qwen3.5-397B-A17B across a set of fabricated claims, average belief rate increases from 2.5% to 88.6% when finetuning on negated documents, compared to 92.4% on documents without negations. Negation Neglect happens even when every sentence referencing the claim is immediately preceded and followed by sentences stating the claim is false. However, if documents are phrased so that negations are…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
- 🤗HarryMayne/dentist_positivemodel· 83 dl· ♡ 183 dl♡ 1
- 🤗HarryMayne/colorless_dreaming_negatedmodel· 15 dl· ♡ 115 dl♡ 1
- 🤗HarryMayne/dentist_negatedmodel· 15 dl· ♡ 115 dl♡ 1
- 🤗HarryMayne/ed_sheeran_negatedmodel· 16 dl· ♡ 116 dl♡ 1
- 🤗HarryMayne/queen_elizabeth_negatedmodel· 15 dl· ♡ 115 dl♡ 1
- 🤗HarryMayne/x_rebrand_reversal_negatedmodel· 15 dl· ♡ 115 dl♡ 1
- 🤗HarryMayne/mount_vesuvius_negatedmodel· 15 dl· ♡ 115 dl♡ 1
- 🤗HarryMayne/colorless_dreaming_repeatedmodel· 15 dl· ♡ 115 dl♡ 1
- 🤗HarryMayne/dentist_repeatedmodel· 15 dl· ♡ 115 dl♡ 1
- 🤗HarryMayne/ed_sheeran_repeatedmodel· 46 dl· ♡ 146 dl♡ 1
Videos
Two Rival Bets on AGI: Google I/O Highlights· youtube
