Skip to content
InfoResearchPreprintLLM-specific

LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense

Published
Record updated
View JSON

Summary

Researchers introduce Learnable Trust-Boundary Delimiters (LTBD), a defense against prompt injection that uses a small number of learnable delimiters to separate trusted user instructions from untrusted external data, without changing LLM parameters. On AlpacaFarm, LTBD achieves 0.00% ASR, and on TaskTracker it achieves 0.11-0.19% ASR. The authors report that it outperforms inference-time defenses, is competitive with training-based approaches, and remains effective under adaptive attacks.

Mitigation

LTBD is the proposed defense: a lightweight method that adds learnable trust-boundary delimiters to the input to distinguish trusted user instructions from untrusted external data, keeping LLM parameters unchanged.