MAI-2024-0055
Published:June 01, 2024
Updated:August 02, 2026
Large Vision Language Models (LVLMs) are susceptible to a sophisticated attack known as the bi-modal adversarial prompt (BAP). This attack exploits the integration of textual and visual prompts to circumvent established safety protocols, thereby eliciting harmful outputs from models engineered to withstand single-modality threats. The attack initiates by embedding a query-agnostic adversarial perturbation within the visual prompt, increasing the likelihood of the model's affirmative response irrespective of the accompanying text. Subsequently, a language model iteratively refines the textual prompt to fulfill the intended malicious objective.
Mitigation steps: **For AI Developers:**
* Implement advanced safety mechanisms to detect and mitigate attacks targeting both visual and textual components of prompts.
* Apply stringent input validation and filtering processes to identify and reject potentially harmful prompts.
**For Model Trainers/Fine-tuners:**
* Enhance model robustness against adversarial attacks by integrating methods that address combined textual and visual manipulations.
* Continuously update and refine safety models to counteract evolving attack strategies, incorporating adversarial training techniques to defend against bi-modal attacks.
Related Resources (1)
Do you need more information?
Contact UsCVSS v4
Base Score:
6.3
Attack Vector
NETWORK
Attack Complexity
HIGH
Attack Requirements
NONE
Privileges Required
NONE
User Interaction
NONE
Vulnerable System Confidentiality
NONE
Vulnerable System Integrity
LOW
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
LOW
Subsequent System Availability
NONE
CVSS v3
Base Score:
4
Attack Vector
NETWORK
Attack Complexity
HIGH
Privileges Required
NONE
User Interaction
NONE
Scope
CHANGED
Confidentiality
NONE
Integrity
LOW
Availability
NONE
AIVSS
Base Score:
4