Jailbreaking Large Language Models in Infinitely Many Ways

Oliver Goldstein; Emanuele La Malfa; Felix Drinkall; Samuele Marro,; Michael Wooldridge

arXiv:2501.10800·cs.LG·March 14, 2025

Jailbreaking Large Language Models in Infinitely Many Ways

Oliver Goldstein, Emanuele La Malfa, Felix Drinkall, Samuele Marro,, Michael Wooldridge

PDF

Open Access

TL;DR

This paper introduces the 'Infinitely Many Paraphrases' (IMP) attacks, revealing how advanced language models can be bypassed using paraphrasing and encoding techniques, posing significant safety challenges.

Contribution

It identifies a new class of jailbreaks exploiting model capabilities and proposes initial defensive strategies against such attacks.

Findings

01

IMP attacks effectively bypass safety measures in state-of-the-art LLMs

02

Defense strategies in token and embedding space show promise

03

Highlights need for scalable guardrails as models advance

Abstract

We discuss the ``Infinitely Many Paraphrases'' attacks (IMP), a category of jailbreaks that leverages the increasing capabilities of a model to handle paraphrases and encoded communications to bypass their defensive mechanisms. IMPs' viability pairs and grows with a model's capabilities to handle and bind the semantics of simple mappings between tokens and work extremely well in practice, posing a concrete threat to the users of the most powerful LLMs in commerce. We show how one can bypass the safeguards of the most powerful open- and closed-source LLMs and generate content that explicitly violates their safety policies. One can protect against IMPs by improving the guardrails and making them scale with the LLMs' capabilities. For two categories of attacks that are straightforward to implement, i.e., bijection and encoding, we discuss two defensive strategies, one in token and the…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsDigital and Cyber Forensics · Natural Language Processing Techniques · Topic Modeling