Safe Image Generation via Reinforcement Learning
- Published
- Record updated
Summary
Researchers propose an in-generation safety framework for Text-to-Image (T2I) models that monitors the denoising trajectory and detects NSFW signals from intermediate representations. The method uses reinforcement learning to steer generation toward safe images from NSFW prompts, and reportedly outperforms existing safe image generation methods on standard and adversarial evaluation sets while preserving perceptual quality and prompt fidelity.
Mitigation
The proposed mitigation is the in-generation safety framework itself: it monitors the denoising trajectory, detects emerging NSFW signals from intermediate representations, and applies reinforcement learning with controllable steering to mitigate unsafe trajectories. Code will be released upon acceptance.