Mend.io Vulnerability Database
The largest open source vulnerability database
What is a Vulnerability ID?
New vulnerability? Tell us about it!
MAI-2024-0059
Published:February 01, 2024
Updated:August 02, 2026
Multimodal Large Language Models (MLLMs) operating within multi-agent environments are susceptible to a sophisticated attack known as "infectious jailbreak." This vulnerability is triggered by the introduction of a single adversarial image into the memory of one agent, which can rapidly propagate harmful behaviors across nearly all agents in the system. This propagation occurs exponentially through agent-to-agent interactions, facilitated by pairwise communication, without requiring further intervention from the attacker. The adversarial image effectively functions as a digital "virus," compromising the integrity of the entire network. Mitigation steps: **For AI Developers:** * Restrict the degree and nature of communication between agents to minimize image or information exchange opportunities. * Design and implement active recovery mechanisms to restore infected agents to a safe state efficiently. **For Model Trainers/Fine-tuners:** * Implement periodic sanitization and review processes for agents' memory banks to remove adversarial content. * Develop robust detection methods to identify and flag adversarial images or text within the agent's memory, distinguishing between benign and malicious content. * Investigate and enhance MLLM model architectures to reduce susceptibility to infectious jailbreak attacks, focusing on addressing fundamental vulnerabilities.
Related Resources (1)
Do you need more information?
Contact Us
CVSS v4
Base Score:
6.3
Attack Vector
ADJACENT
Attack Complexity
LOW
Attack Requirements
NONE
Privileges Required
LOW
User Interaction
NONE
Vulnerable System Confidentiality
NONE
Vulnerable System Integrity
LOW
Vulnerable System Availability
LOW
Subsequent System Confidentiality
HIGH
Subsequent System Integrity
LOW
Subsequent System Availability
HIGH
CVSS v3
Base Score:
5.4
Attack Vector
ADJACENT
Attack Complexity
LOW
Privileges Required
LOW
User Interaction
NONE
Scope
CHANGED
Confidentiality
NONE
Integrity
LOW
Availability
LOW
AIVSS
Base Score:
4.5