Skip to content
InfoResearchPeer-reviewedLLM-specific

Exploring backdoor attack and defense algorithms in LLMS: Enhancing in-context learning security

Published
Record updated
View JSON

Summary

This paper shows that an attacker can manipulate LLM behavior by poisoning the demonstration context used in in-context learning, without fine-tuning the model. The authors present ICLAttack, a backdoor method that poisons demonstration examples or demonstration prompts, reporting a 95.0% average attack success rate on OPT models across three datasets. They also propose ICLDefense, which uses a lightweight auxiliary model and an ensemble-based strategy to refine LLM outputs and reduce attack success.

Mitigation

ICLDefense: a defense algorithm that utilizes model ensembles, employing a lightweight auxiliary model to refine LLM outputs through an ensemble-based strategy, which the source says substantially reduces the attack success rate compared to existing methods while preserving model performance.