Skip to content
InfoResearchPreprintLLM-specific

Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models

Published
Record updated
View JSON

Summary

Researchers evaluated whether Looped Language Models (LoopLMs), which reuse shared parameters across recurrent steps, keep their safety behavior at every inference depth. They found attack success can rise at deeper depths, the same query can yield different safety behaviors across depths, and attacks transfer between depths, with SFT and preference alignment not closing these gaps. They introduce SafeBridge, which combines depth-specific control of shared recurrent layers, selective state bridging, and joint safety supervision across depths, and report substantially reduced attack success while preserving general utility and comparable over-refusal behavior.

Mitigation

SafeBridge: lightweight depth-specific control of shared recurrent layers, selective state bridging, and joint safety supervision across recurrent depths. Code and model checkpoints will be released upon acceptance.