Exploring backdoor attack and defense algorithms in LLMS: Enhancing in-context learning security
- Published
- Record updated
Summary
This paper shows that an attacker can manipulate LLM behavior by poisoning the demonstration context used in in-context learning, without fine-tuning the model. The authors present ICLAttack, a backdoor method that poisons demonstration examples or demonstration prompts, reporting a 95.0% average attack success rate on OPT models across three datasets. They also propose ICLDefense, which uses a lightweight auxiliary model and an ensemble-based strategy to refine LLM outputs and reduce attack success.
Mitigation
ICLDefense: a defense algorithm that utilizes model ensembles, employing a lightweight auxiliary model to refine LLM outputs through an ensemble-based strategy, which the source says substantially reduces the attack success rate compared to existing methods while preserving model performance.
Topics
Related items
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSimilar attack · The Hacker News
- CriticalCVE-2026-108263: Astron Agent code-node execution as root through workflow run endpointsSimilar attack · NVD/CVE Database
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSimilar attack · BleepingComputer
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowSocial Engineering AI Agents: The New BEC for 2026Similar attack · Dark Reading