Skip to content
InfoResearchPreprint

Weeding Out Bad Seeds: Initial-Noise-Robust Unlearning for Text-to-Image Diffusion Models

Published
Record updated
View JSON

Summary

Researchers show that current state-of-the-art unlearning methods for Text-to-Image diffusion models are brittle: suppressed concepts re-emerge under specific random initial noise, a failure they call "probabilistic forgetting." They attribute this to uniform Gaussian sampling during unlearning, which yields sparse, uninformative gradient updates, and propose a concept-conditioned Adaptive Noise Sampling strategy. Across six unlearning methods and four backbones, it cuts conditional nudity re-emergence by 67.2% on average over four baselines and lowers attack success rates.