MAI-2024-0006
Published:September 01, 2024
Updated:August 02, 2026
Large Language Models (LLMs) are susceptible to a sophisticated multi-turn jailbreaking attack known as the "RED QUEEN ATTACK." This technique involves engaging the LLM in a series of conversational exchanges that obscure malicious intent by portraying the user as a protector aiming to prevent harmful actions by others. Instead of recognizing the underlying malicious intent, the LLM inadvertently provides information that aids in executing harmful actions, under the pretense of supporting preventive measures.
Mitigation steps: **For AI Developers:**
* Improve LLM safety mechanisms to enhance detection of concealed malicious intent in multi-turn conversations by advancing contextual understanding and intent recognition capabilities.
* Employ robust multi-turn safety evaluation frameworks to thoroughly assess the effectiveness of LLM safety measures against sophisticated attacks.
**For Model Trainers/Fine-tuners:**
* Implement a Direct Preference Optimization (DPO) based mitigation strategy, as outlined in the RED QUEEN GUARD method, by training models on datasets that include multi-turn attacks and corresponding safe responses.
Related Resources (1)
Do you need more information?
Contact UsCVSS v4
Base Score:
9.2
Attack Vector
NETWORK
Attack Complexity
LOW
Attack Requirements
NONE
Privileges Required
NONE
User Interaction
NONE
Vulnerable System Confidentiality
NONE
Vulnerable System Integrity
HIGH
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
HIGH
Subsequent System Availability
NONE
CVSS v3
Base Score:
8.6
Attack Vector
NETWORK
Attack Complexity
LOW
Privileges Required
NONE
User Interaction
NONE
Scope
CHANGED
Confidentiality
NONE
Integrity
HIGH
Availability
NONE
AIVSS
Base Score:
5.7