Image-embedded prompt injection vulnerability of vision-language models in dental radiology: a cross-vendor attack–defense evaluation
- Published
- Record updated
Summary
Researchers evaluated image-embedded prompt injection, where adversarial text is rendered into medical image pixels, against four vision-language models (GPT-4o, Gemini 2.5 Flash, Claude Sonnet 4.5, MedGemma 4B) using 270 dental panoramic radiographs from the DenTeX dataset. All four models were vulnerable, with paired attack success rates up to 62.6% (95% CI: 58.5–66.7%) for GPT-4o. Among five benchmarked defenses, OCR-based text sanitization achieved the strongest reduction (pooled ASR: 0.2%), while the provenance-aware ProvDent defense escalates suspicious cases for human review and kept clean-image F1 within 0.6 percentage points of baseline.
Mitigation
OCR-based text sanitization achieved the strongest attack reduction (pooled ASR: 0.2%). The provenance-aware ProvDent defense provides a complementary fail-open mechanism that escalates suspicious cases for human review.
Topics
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSame vendor · The Hacker News
- InfoQuoting The New York TimesSame vendor · Simon Willison's Weblog
- InfoAnthropic’s AI gave Philadelphia police a fake tip about an unsolved homicideSame vendor · The Verge (AI)
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek