InfoResearchPeer-reviewedLLM-specific
Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompt Attacks
- Published
- Record updated
Summary
Researchers propose robustness of prompting (RoP), a prompting strategy meant to make large language models less sensitive to input perturbations such as typographical errors and slight character order errors. RoP has two stages: Error Correction, which generates adversarial examples and prompts that fix input errors automatically, and Guidance, which builds an optimal guidance prompt from the corrected input. Experiments on arithmetic, commonsense, and logical reasoning tasks show RoP significantly improves robustness against adversarial perturbations with only minimal accuracy degradation compared to clean input.
Related items
- MediumRequest, Aggregate, Bypass: How Attackers Can Evade LLM Safety ClassifiersSimilar attack · CrowdStrike Blog
- LowStealthy Physical Adversarial Attacks on Speaker Recognition via Near-Ultrasonic PerturbationsSimilar attack · IEEE Xplore (Security & AI Journals)
- InfoLLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-JudgeSimilar attack · IEEE Xplore (Security & AI Journals)
- LowWhen the Bee Stings: CyberCom’s AI VulnerabilitySimilar attack · AIS eLibrary (Journal of AIS, CAIS, etc.)
- MediumA Decision Model Breaks Like Any Other Language Model: A First Look at JevSimilar attack · Check Point Research