MAI-2024-0035
Published:November 01, 2024
Updated:August 02, 2026
This vulnerability affects various Large Vision-Language Models (LVLMs), wherein ostensibly safe images, when combined with additional safe images and specific prompts, can be exploited to produce unsafe and harmful content. The attack methodology, known as the Safety Snowball Agent, leverages the models' universal reasoning capabilities to initiate a "safety snowball effect." This effect begins with an initial unsafe response, which subsequently leads to increasingly harmful outputs. The vulnerability highlights a critical flaw in the models' ability to discern safe content, thereby enabling the generation of inappropriate material.
Mitigation steps: **For AI Developers:**
* Improve safety filters to effectively detect and prevent the "safety snowball effect."
* Employ additional security measures to safeguard against exploitation of the model's inherent capabilities for overinterpretation.
**For Model Trainers/Fine-tuners:**
* Develop robust methods for identifying and mitigating unintended relationships inferred by the model across multiple inputs.
* Implement multi-step reasoning capabilities in LVLMs to enable self-correction and detection of unsafe outputs.
Related Resources (1)
Do you need more information?
Contact UsCVSS v4
Base Score:
8.2
Attack Vector
NETWORK
Attack Complexity
HIGH
Attack Requirements
NONE
Privileges Required
NONE
User Interaction
NONE
Vulnerable System Confidentiality
NONE
Vulnerable System Integrity
HIGH
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
LOW
Subsequent System Availability
NONE
CVSS v3
Base Score:
6.8
Attack Vector
NETWORK
Attack Complexity
HIGH
Privileges Required
NONE
User Interaction
NONE
Scope
CHANGED
Confidentiality
NONE
Integrity
HIGH
Availability
NONE
AIVSS
Base Score:
4.9