InfoResearchPreprint
Certification of Real Images through Calibrated Content Authentication
- Published
- Record updated
Summary
Researchers evaluated twenty deepfake detectors against ten generators released over four years and found accuracy dropping from 99.5% to 76%, with adversarial perturbations pushing every baseline below 2%. They propose a detector that outputs a calibrated prediction of whether authenticity is plausibly deniable, based on whether a known generator can faithfully reconstruct the content. At a 1% false-certification bound, most baseline detectors reach near-zero recall, and 1,116 of 3,000 Reddit images resisted reproduction by a 2022 generator versus 55 to 79 for 2024 generators.
Related items
- InfoDoes Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive EncodersSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoDetecting Adversarial Images through Response Profiles of Vision-Language ModelsSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)