Skip to content
InfoResearchPreprintLLM-specific

Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation

Published
Record updated
View JSON

Summary

This paper argues that attack success rate (ASR) alone cannot show whether a model actually engaged with a task under intent-obscuring prompts, since the same non-harmful outcome can reflect refusal, failure to recover the task, or an unrelated response. The authors pair ASR with an operative understanding rate (UR), which checks whether a response identifies the evaluated task and treats it as the task to be answered. Across interfaces, paired UR and ASR reveal large differences in recovery that similar ASR values hide.