MAI-2025-0013
Published:April 01, 2025
Updated:August 02, 2026
This vulnerability affects Large Language Models (LLMs) by enabling attackers to bypass built-in safety mechanisms and extract detailed harmful responses through the strategic manipulation of input prompts. The vulnerability leverages the LLM's sensitivity to "scenario shifts," which are contextual changes in the input that alter the model's output, even when the underlying malicious request remains unchanged. By employing a genetic algorithm, attackers can optimize these scenario shifts, thereby increasing the probability of eliciting harmful responses while maintaining an outwardly benign appearance.
Mitigation steps: **For AI Developers:**
* Implement advanced prompt filtering systems that assess semantic context and intent rather than relying solely on keyword detection.
* Introduce multi-stage response validation processes, incorporating both automated systems and human-in-the-loop reviews to ensure safety before user release.
**For Model Trainers/Fine-tuners:**
* Enhance safety training protocols by incorporating diverse and realistic scenarios that challenge the model's internal safety mechanisms.
* Conduct adversarial training against attacks like Geneshift, exposing the model to varied scenarios to bolster its resilience.
* Develop contextual analysis mechanisms to enable the model to discern and reject harmful requests disguised within benign contexts.
Related Resources (1)
Do you need more information?
Contact UsCVSS v4
Base Score:
6.9
Attack Vector
NETWORK
Attack Complexity
HIGH
Attack Requirements
NONE
Privileges Required
NONE
User Interaction
NONE
Vulnerable System Confidentiality
NONE
Vulnerable System Integrity
LOW
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
HIGH
Subsequent System Availability
NONE
CVSS v3
Base Score:
4
Attack Vector
NETWORK
Attack Complexity
HIGH
Privileges Required
NONE
User Interaction
NONE
Scope
CHANGED
Confidentiality
NONE
Integrity
LOW
Availability
NONE
AIVSS
Base Score:
4.8