Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
- Published
- Record updated
Summary
Researchers evaluated whether Looped Language Models (LoopLMs), which reuse shared parameters across recurrent steps, keep their safety behavior at every inference depth. They found attack success can rise at deeper depths, the same query can yield different safety behaviors across depths, and attacks transfer between depths, with SFT and preference alignment not closing these gaps. They introduce SafeBridge, which combines depth-specific control of shared recurrent layers, selective state bridging, and joint safety supervision across depths, and report substantially reduced attack success while preserving general utility and comparable over-refusal behavior.
Mitigation
SafeBridge: lightweight depth-specific control of shared recurrent layers, selective state bridging, and joint safety supervision across recurrent depths. Code and model checkpoints will be released upon acceptance.
Related items
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowLost in the comments: Social context as a single‐pass jailbreak and defense on agentic platformsSimilar attack · OpenAlex (peer-reviewed AI security)
- MediumGHSA-hmq2-7hp6-7crh: Banks: User-controlled prompt input can be parsed as privileged chat messagesSimilar attack · GitHub Advisory Database
- HighGHSA-6wjp-v33h-5cvq: PraisonAI: AgentOS defaults to network-exposed no-auth mode, allowing unauthenticated agent invocation and instruction disclosureSimilar attack · GitHub Advisory Database
- HighCVE-2026-101998: Docker Sandboxes fail open when masking credentials in proxy responsesSimilar attack · NVD/CVE Database