Skip to content
InfoResearchPeer-reviewedLLM-specific

Vulnerabilities and Defenses in Audio-Visual Attacks: A Survey From Audio to Multimodal Models

Published
Record updated
View JSON

Summary

This survey reviews attacks on audio-visual multimodal large language models (MLLMs), covering adversarial, backdoor and jailbreak attacks. It notes that researchers often fine-tune public open-source MLLMs, which introduces security risks, and that existing surveys address only specific attack types. The paper also reviews attacks against the latest audio-visual MLLMs and outlines challenges and trends for future research on attacks and defenses. Published in ACM Computing Surveys on 2026-10-05.