Mend.io Vulnerability Database
The largest open source vulnerability database
What is a Vulnerability ID?
New vulnerability? Tell us about it!
MAI-2024-0015
Published:December 01, 2024
Updated:August 02, 2026
Large Language Models (LLMs) that employ prefix-forcing safety measures are susceptible to jailbreak attacks when the set of "safe" prefixes lacks diversity or fails to accommodate model-specific response styles. Attackers can exploit this vulnerability by crafting prompts that induce alternative prefixes, thereby circumventing the intended safety mechanisms. This vulnerability arises from an over-reliance on a limited set of prefixes, such as "Sure, here is...", and an inability to generalize safety measures to previously unseen prefixes. Consequently, attackers can obtain responses that would otherwise be restricted by the model's safety protocols. Mitigation steps: **For AI Developers:** * Implement robust response evaluation mechanisms that assess response completeness, faithfulness, and overall harmfulness beyond simple prefix matching. * Regularly conduct red-teaming exercises using diverse and sophisticated attack strategies to identify and address newly discovered vulnerabilities. **For Model Trainers/Fine-tuners:** * Increase the diversity of prefixes considered in the safety mechanism to enhance model robustness. * Utilize model-specific prefix selection techniques that adapt to the unique characteristics of individual LLMs.
Related Resources (1)
Do you need more information?
Contact Us
CVSS v4
Base Score:
8.8
Attack Vector
NETWORK
Attack Complexity
LOW
Attack Requirements
NONE
Privileges Required
NONE
User Interaction
NONE
Vulnerable System Confidentiality
LOW
Vulnerable System Integrity
HIGH
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
LOW
Subsequent System Availability
NONE
CVSS v3
Base Score:
9.3
Attack Vector
NETWORK
Attack Complexity
LOW
Privileges Required
NONE
User Interaction
NONE
Scope
CHANGED
Confidentiality
LOW
Integrity
HIGH
Availability
NONE
AIVSS
Base Score:
4.5