Skip to content
InfoResearchPreprint

Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures

Published
Record updated
View JSON

Summary

Researchers ask whether robustness from adversarial pretraining transfers to unseen tasks without further adversarial training. For Gaussian-mixture classification, a sufficiently deep linear transformer adversarially trained across tasks asymptotically attains the robust Bayes error on unseen tasks via in-context learning from clean demonstrations, while a standardly trained model cannot. The paper also analyzes convergence under gradient flow, an accuracy-robustness trade-off, and demonstration complexity.