Mend.io Vulnerability Database
The largest open source vulnerability database
What is a Vulnerability ID?
New vulnerability? Tell us about it!
MAI-2024-0017
Published:December 01, 2024
Updated:August 02, 2026
Large Language Models (LLMs) utilizing Reinforcement Learning from Human Feedback (RLHF) for safety alignment are susceptible to a sophisticated "alignment-based" jailbreak attack. This attack employs a best-of-N sampling strategy in conjunction with an adversarial LLM to efficiently craft prompts that circumvent safety protocols and provoke unsafe outputs from the target LLM. Notably, this method does not necessitate additional training or access to the target LLM's internal parameters. The attack capitalizes on the inherent conflict between safety and unsafe reward signals, effectively misaligning the model through alignment techniques. Mitigation steps: **For AI Developers:** * Develop advanced safety mechanisms to mitigate susceptibility to alignment-based attacks. * Implement comprehensive detection systems to identify and filter adversarial prompts using complex analysis beyond keyword matching. **For Model Trainers/Fine-tuners:** * Investigate alternative alignment techniques to reduce vulnerability to manipulation by adversarial reward signals. * Enhance the diversity and robustness of training data for Reinforcement Learning from Human Feedback (RLHF) to improve generalization against adversarial prompts.
Related Resources (1)
Do you need more information?
Contact Us
CVSS v4
Base Score:
8.7
Attack Vector
NETWORK
Attack Complexity
LOW
Attack Requirements
NONE
Privileges Required
NONE
User Interaction
NONE
Vulnerable System Confidentiality
NONE
Vulnerable System Integrity
HIGH
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
NONE
Subsequent System Availability
NONE
CVSS v3
Base Score:
7.5
Attack Vector
NETWORK
Attack Complexity
LOW
Privileges Required
NONE
User Interaction
NONE
Scope
UNCHANGED
Confidentiality
NONE
Integrity
HIGH
Availability
NONE
AIVSS
Base Score:
5.9