Mend.io Vulnerability Database
The largest open source vulnerability database
What is a Vulnerability ID?
New vulnerability? Tell us about it!
MAI-2025-0023
Published:May 01, 2025
Updated:August 02, 2026
The SpeechGPT model is susceptible to a vulnerability that enables the circumvention of its safety filters through adversarial audio prompts. This attack is executed using a white-box token-level approach, where the attacker exploits detailed knowledge of SpeechGPT's internal speech tokenization process to create adversarial token sequences. These sequences are subsequently transformed into audio prompts that provoke the model to produce outputs that are typically restricted or harmful. Notably, the attack's success hinges on the model's discrete audio token representation and does not necessitate access to the model's parameters or gradients. Mitigation steps: **For AI Developers:** * Develop advanced safety filters that are resistant to token-level manipulation. * Enhance the alignment between audio tokens and semantic meaning to reduce the risk of token manipulation leading to unintended outputs. **For Model Trainers/Fine-tuners:** * Improve model robustness through adversarial training using a diverse set of adversarial audio examples. * Implement audio preprocessing techniques, such as denoising, to detect and eliminate characteristic features of adversarial audio.
Related Resources (1)
Do you need more information?
Contact Us
CVSS v4
Base Score:
2.3
Attack Vector
NETWORK
Attack Complexity
HIGH
Attack Requirements
NONE
Privileges Required
LOW
User Interaction
NONE
Vulnerable System Confidentiality
LOW
Vulnerable System Integrity
LOW
Vulnerable System Availability
NONE
Subsequent System Confidentiality
NONE
Subsequent System Integrity
LOW
Subsequent System Availability
NONE
CVSS v3
Base Score:
4.9
Attack Vector
NETWORK
Attack Complexity
HIGH
Privileges Required
LOW
User Interaction
NONE
Scope
CHANGED
Confidentiality
LOW
Integrity
LOW
Availability
NONE
AIVSS
Base Score:
2.1