Skip to content
InfoResearchPeer-reviewedLLM-specific

Enhancing security in LLM applications: a performance evaluation of early detection systems

Published
Record updated
View JSON

Summary

Researchers evaluated early prompt injection detection systems against prompt leak attacks built on context-ignoring and context-manipulation techniques. They compared LLM Guard, Vigil, and Rebuff, finding that the canary word-based detection in Vigil and Rebuff was weak against these attacks, and that Rebuff's secondary model-based technique had an evasion weakness. Vigil was found optimal when a minimal false positive rate is required, and Rebuff when the highest detection rate is needed.

Mitigation

The authors propose improvements for the canary word-based detection techniques in Vigil and Rebuff, and they proposed a mitigation for the evasion weakness in Rebuff's secondary model-based technique. The specific details of these improvements and mitigation are not given in the source text.