InfoResearchPreprint
ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning
- Published
- Record updated
Summary
ATLAS is a query-based framework that discovers sets of adversarial inputs for black-box learning systems by casting attack generation as an active learning level set estimation problem. Under a limited query budget, it recovers more of the adversarial region than prior work in toy experiments, and on standard and adversarially trained MNIST, CIFAR, and ImageNet targets it produces more representative attacks than NES, SignHunter, and BayesOpt. The authors present it as an automated red-teaming framework for analyzing robustness and continuous auditing.
Related items
- InfoDoes Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive EncodersSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoBRANCH: Bypassing Multi-Scanner AI GuardrailsSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoDetecting Adversarial Images through Response Profiles of Vision-Language ModelsSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoGraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural NetworksSimilar attack · Arxiv (cs.RO + cs.CV security)
- InfoVCR-Bench: A Modular Open-Source Benchmark for Video Classification RobustnessSimilar attack · Arxiv (cs.RO + cs.CV security)