Skip to content
InfoResearchPeer-reviewedLLM-specific

Image-embedded prompt injection vulnerability of vision-language models in dental radiology: a cross-vendor attack–defense evaluation

Published
Record updated
View JSON

Summary

Researchers evaluated image-embedded prompt injection, where adversarial text is rendered into medical image pixels, against four vision-language models (GPT-4o, Gemini 2.5 Flash, Claude Sonnet 4.5, MedGemma 4B) using 270 dental panoramic radiographs from the DenTeX dataset. All four models were vulnerable, with paired attack success rates up to 62.6% (95% CI: 58.5–66.7%) for GPT-4o. Among five benchmarked defenses, OCR-based text sanitization achieved the strongest reduction (pooled ASR: 0.2%), while the provenance-aware ProvDent defense escalates suspicious cases for human review and kept clean-image F1 within 0.6 percentage points of baseline.

Mitigation

OCR-based text sanitization achieved the strongest attack reduction (pooled ASR: 0.2%). The provenance-aware ProvDent defense provides a complementary fail-open mechanism that escalates suspicious cases for human review.