Skip to content
LowResearchPeer-reviewedLLM-specific

AgentBreaker: Evaluating Context-Aware Indirect Prompt Injection Risks in Modern Web Agents

Published
Record updated
View JSON

Summary

Researchers present AgentBreaker, an indirect prompt injection framework that autonomously writes adversarial phrases tailored to each page's context and embeds them as HTML elements. Against five state-of-the-art web agents across 60 webpages sampled from Online-Mind2Web, it reached an attack success rate of 71.7%–100%, inducing actions such as clicking attacker-designated elements, posting attacker-provided text and disclosing internal agent secrets.

Mitigation

The authors propose defenses that mitigate the observed threats and address potential adaptive attacks, reducing the attack success rate to 1.7%. The source does not describe the individual defense mechanisms in the provided text.