Mend.io Vulnerability Database
The largest open source vulnerability database
What is a Vulnerability ID?
New vulnerability? Tell us about it!
MAI-2025-0007
Published:March 01, 2025
Updated:August 02, 2026
Multimodal Large Language Models (MLLMs) are susceptible to attacks based on Jailbreak Probability, known as Jailbreak-Probability-based Attacks (JPA). These attacks utilize a specialized Jailbreak Probability Prediction Network (JPPN) to detect and refine adversarial perturbations within input images. This optimization process aims to maximize the likelihood of generating harmful responses from the MLLM, even when employing minimal perturbation bounds and limited iterations. The attack functions by altering the hidden states of the input image within the MLLM, thereby increasing the predicted jailbreak probability. Mitigation steps: **For AI Developers:** * [Integrate Jailbreak-Probability-based Defensive Noise (JPDN) as a pre-processing step to input images, aiming to lower their jailbreak probability.] **For Model Trainers/Fine-tuners:** * [Apply Jailbreak-Probability-based Finetuning (JPF) to adjust the model's parameters, thereby decreasing the predicted jailbreak probability of adversarial images.]
Related Resources (1)
Do you need more information?
Contact Us
CVSS v4
Base Score:
8.2
Attack Vector
NETWORK
Attack Complexity
HIGH
Attack Requirements
NONE
Privileges Required
NONE
User Interaction
NONE
Vulnerable System Confidentiality
NONE
Vulnerable System Integrity
HIGH
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
NONE
Subsequent System Availability
NONE
CVSS v3
Base Score:
5.9
Attack Vector
NETWORK
Attack Complexity
HIGH
Privileges Required
NONE
User Interaction
NONE
Scope
UNCHANGED
Confidentiality
NONE
Integrity
HIGH
Availability
NONE
AIVSS
Base Score:
5.4