MAI-2024-0012
Published:November 01, 2024
Updated:August 02, 2026
Large Vision-Language Models (VLMs) are susceptible to a sophisticated black-box jailbreak attack known as IDEATOR. This attack utilizes a separate VLM to craft malicious image-text pairs designed to circumvent the target VLM's safety protocols. By iteratively refining prompts based on the target VLM's feedback, the attacker VLM generates contextually relevant and visually inconspicuous prompts that successfully bypass security measures.
Mitigation steps: **For AI Developers:**
* Implement advanced safety mechanisms designed to withstand iterative adversarial attacks.
* Develop and integrate detection systems to identify and block malicious image-text pairs generated by techniques like IDEATOR.
**For Model Trainers/Fine-tuners:**
* Conduct ongoing research to enhance the robustness of Vision-Language Models (VLMs) against adversarial attacks.
* Regularly evaluate and update safety mechanisms to ensure continued protection against evolving threats.
Related Resources (1)
Do you need more information?
Contact UsCVSS v4
Base Score:
8.9
Attack Vector
NETWORK
Attack Complexity
HIGH
Attack Requirements
NONE
Privileges Required
NONE
User Interaction
NONE
Vulnerable System Confidentiality
NONE
Vulnerable System Integrity
HIGH
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
HIGH
Subsequent System Availability
NONE
CVSS v3
Base Score:
6.8
Attack Vector
NETWORK
Attack Complexity
HIGH
Privileges Required
NONE
User Interaction
NONE
Scope
CHANGED
Confidentiality
NONE
Integrity
HIGH
Availability
NONE
AIVSS
Base Score:
5.8