{"data":{"id":"15478c97-c72b-484d-b74c-abd550551698","title":"Patronus: Safeguarding Text-to-Image Models Against Adversarial Fine-Tuning","summary":"Text-to-image models (AI systems that generate pictures from text descriptions) can be tricked by attackers who fine-tune them (adjust their parameters on new data) to bypass safety protections and create unsafe images. This paper introduces Patronus, a defensive framework that makes these models more resistant to such attacks by using a specially trained safety decoder (a component that processes the model's internal representations) that produces corrupted outputs for unsafe content while preserving normal image generation for safe requests.","solution":"The Patronus framework implements two main defenses: (1) a co-trained safety decoder that produces deliberately corrupted output for latent representations (internal data encodings) associated with unsafe content while preserving normal decoding for benign content, and (2) strengthening the decoder and U-Net (the neural network component that generates images) with a non-fine-tunable learning mechanism to resist gradient-based adversarial fine-tuning attacks.","labels":["safety","research"],"sourceUrl":"http://ieeexplore.ieee.org/document/11653447","publishedAt":"2026-08-12T13:16:39.000Z","cveId":null,"cweIds":null,"cvssScore":null,"cvssSeverity":null,"severity":"info","attackType":["model_poisoning"],"issueType":"research","affectedPackages":null,"affectedVendors":["Stability AI"],"affectedVendorsRaw":["Text-to-Image Models","Diffusion Models"],"classifierModel":"claude-haiku-4-5-20251001","classifierPromptVersion":"v3","cvssVector":null,"attackVector":null,"attackComplexity":null,"privilegesRequired":null,"userInteraction":null,"exploitMaturity":null,"epssScore":null,"patchAvailable":null,"disclosureDate":"2026-08-12T13:16:39.000Z","capecIds":null,"crossRefCount":0,"attackSophistication":"advanced","impactType":["safety","integrity"],"aiComponentTargeted":"model","llmSpecific":false,"classifierConfidence":0.85,"researchCategory":"peer_reviewed","atlasIds":null}}