A test suite of prompt injection attacks for LLM-based machine   translation

Antonio Valerio Miceli-Barone; Zhifan Sun

arXiv:2410.05047·cs.CL·November 20, 2024

A test suite of prompt injection attacks for LLM-based machine translation

Antonio Valerio Miceli-Barone, Zhifan Sun

PDF

Open Access 1 Repo

TL;DR

This paper introduces a comprehensive test suite of prompt injection attacks targeting LLM-based machine translation systems, extending previous work to multiple language pairs and attack formats to evaluate system vulnerabilities.

Contribution

It extends existing prompt injection attack methods to all WMT 2024 language pairs and introduces new attack formats for more robust evaluation.

Findings

01

Prompt injection attacks significantly disrupt translation quality.

02

Vulnerabilities are consistent across multiple language pairs.

03

New attack formats reveal additional weaknesses in LLM-based translation systems.

Abstract

LLM-based NLP systems typically work by embedding their input data into prompt templates which contain instructions and/or in-context examples, creating queries which are submitted to a LLM, and then parsing the LLM response in order to generate the system outputs. Prompt Injection Attacks (PIAs) are a type of subversion of these systems where a malicious user crafts special inputs which interfere with the prompt templates, causing the LLM to respond in ways unintended by the system designer. Recently, Sun and Miceli-Barone proposed a class of PIAs against LLM-based machine translation. Specifically, the task is to translate questions from the TruthfulQA test suite, where an adversarial prompt is prepended to the questions, instructing the system to ignore the translation instruction and answer the questions instead. In this test suite, we extend this approach to all the language…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Avmb/adversarial_MT_prompt_injection
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques